UNmiss Blog

Robots.txt Now Shows Who Signed an AI Deal

We read 49 robots.txt files. Seven publishers block Anthropic and allow OpenAI, and all seven have an OpenAI licensing deal. The most-blocked crawler is a non-profit.

Every website has a small file called robots.txt. For thirty years it was a piece of housekeeping: here are the folders I would rather you did not crawl. Nobody read it for fun.

It has quietly become something else.

We opened the robots.txt of 49 well-known websites and checked which AI crawlers each one turns away. What came back was not a technical pattern. It was a map of who has signed a licensing deal.

Seven major publishers block Anthropic's crawler and let OpenAI's walk straight through. Every single one of them has a publicly reported content deal with OpenAI.

That file is not housekeeping any more. It is a contract, published at the root of the domain, where anyone can read it.

Results at a glance

7 of 7
sites blocking Anthropic but not OpenAI have an OpenAI deal
52%
of blockers actually block OpenAI's GPTBot
90%
block Common Crawl, a non-profit research crawler
The short version

Publishers block AI companies they are fighting, and stop blocking the ones that pay them. You can see which is which without reading a single press release.

Meanwhile most blocklists are copied from 2023 and stop the wrong crawlers.

Who blocks anything at all

First, the easy part. Of the 49 sites we checked, 21 block at least one AI crawler — and which side of that line you fall on depends almost entirely on what your business sells.

If your revenue comes from people reading your pages, an AI that answers without sending the reader is a competitor. If your revenue comes from people using your product, an AI that reads your documentation is a salesperson.

Same file, opposite incentive.

The list nobody blocks consistently

Now look at what the 21 blockers actually block, and the tidy story falls apart.

Common Crawl has been quietly archiving the web since 2011, long before anyone worried about chatbots. It is now the single most blocked crawler on this list.

Meanwhile OAI-SearchBot — the crawler that feeds ChatGPT's search results — is blocked by only a third of the sites that went to the trouble of blocking anything.

Some of that is simple staleness. A lot of these lists were written during the 2023 scraping panic, pasted from a blog post, and never revisited. But staleness does not explain the next part.

Seven publishers, one conspicuous gap

Seven sites in our sample block Anthropic's ClaudeBot and do not block OpenAI's GPTBot.

The Guardian. Wired. The Verge. The Washington Post. The Atlantic. Business Insider. Investopedia.

Here is what Wired's file looks like. Read the list of crawlers being turned away, then notice who is missing from it.

An excerpt of wired.com's robots.txt file listing blocked user agents including Google-Extended, Applebot-Extended, Amazonbot, meta-externalagent, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, cohere-ai, CCBot, Bytespider, Diffbot, DuckAssistBot, MistralAI-User and Timpibot.
What to notice: Google, Apple, Amazon, Meta, Anthropic, Perplexity, Cohere, Mistral and ByteDance are all named. OpenAI is not mentioned anywhere in the file.

Anthropic is blocked three times over. Google, Apple, Amazon, Meta, Perplexity, Cohere, Mistral and ByteDance are all turned away. OpenAI's crawlers do not appear in the file at all.

Then we checked all seven names against public reporting on AI licensing agreements.

All seven have a publicly reported content deal with OpenAI. The Guardian, Condé Nast — which owns Wired — Vox Media, which owns The Verge, The Washington Post, The Atlantic, Axel Springer, which owns Business Insider, and Dotdash Meredith, which owns Investopedia.

Seven out of seven. That is not a coincidence you can wave away.

The other group is suing

Look at the sites that block everything instead, and the pattern holds from the opposite direction.

The New York Times blocks all 17 crawlers we tested, OpenAI's included. So does CNN. The BBC blocks 15. These are the organisations in litigation with AI companies or publicly refusing to license, and their files say exactly that.

So a publisher's robots.txt now sorts into two shapes. Block everyone, and you are fighting or holding out. Block everyone except one company, and that company is paying you.

Reuters is the third shape and worth a mention: it names a long list of bots in a permissive group and never mentions GPTBot, ClaudeBot, CCBot or Google-Extended at all. Deciding not to decide is also a position.

What this means for your site

You are probably not negotiating with OpenAI. The lesson still applies, in three parts.

1. Stop copying other people's blocklists. That list was written for somebody else's contracts and somebody else's business model. Copying Wired's file means copying Condé Nast's negotiating position, which is unlikely to be yours.

2. Decide what you are actually protecting. If people pay to read your words, an AI answering in your place costs you money. If people pay to use your product, an AI reading your documentation is doing your marketing. Those two situations do not share a robots.txt file.

3. If you do block, block the right names. Training crawlers and answer crawlers are different things with different agents. Blocking GPTBot stops training but leaves OAI-SearchBot free to cite you in ChatGPT — which is what most people actually want. Blocking CCBot mostly annoys a research archive.

And check what your file currently says, because most people have never read theirs.

UNmiss Website Audit report for wired.com showing a site health score of 83 out of 100, zero critical issues, three warnings, eight notices and 110 checks passed, with groups for on-page SEO, technical SEO, speed and mobile.
What to notice: the crawler check reads your robots.txt as a search engine would, so you can see what you are actually telling bots rather than what you meant to.

What we could not measure

We measured correlation, not cause. Seven for seven is a striking pattern, but we cannot see inside anyone's contract. It is possible a licensing agreement never mentions crawler access and the two things simply travel together. We think that is unlikely; we cannot prove it.

Blocking is a request, not a wall. robots.txt has no enforcement of any kind. A well-behaved crawler obeys it. We have no way of knowing who obeys.

49 sites is a sample. We picked recognisable names across six categories, and a different 49 would move the percentages.

Deals change constantly. This is a photograph taken on 22 August 2026. Both the files and the agreements behind them will have moved by the time you read it, which is exactly why the method matters more than the numbers.

Free, no account
See what your robots.txt is telling crawlers

Our Website Audit fetches your site the way a search engine does and reports what it found, including whether your robots.txt is present and what it permits. It is the fastest way to check that the file at your root says what you think it says.

  • Crawlability, robots.txt and sitemap checks in one report
  • Around 140 technical checks, worst issues first
  • 3 checks a day, no signup, no card
Run a free site audit

Frequently asked questions

What exactly did you check?

We fetched robots.txt from 49 well-known domains on 22 August 2026 and parsed it properly — grouping consecutive User-agent lines with the rules that follow them — then recorded which of 18 AI crawlers had their own group containing a site-wide Disallow: /. Counting a mention of a bot name is not enough, because many sites name bots in permissive groups.

How many sites block AI crawlers?

21 of 49, or 42%. By category: publishers 10 of 12, reference sites 5 of 8, retailers 3 of 8, SaaS 3 of 8, developer tools 0 of 8 and SEO tools 0 of 5. Among the sites that block anything, the average blocklist named about 11 crawlers.

Which seven sites block Anthropic but not OpenAI?

The Guardian, Wired, The Verge, The Washington Post, The Atlantic, Business Insider and Investopedia. Each one, or its parent company, has a publicly reported content licensing agreement with OpenAI: Condé Nast owns Wired, Vox Media owns The Verge, Axel Springer owns Business Insider and Dotdash Meredith owns Investopedia.

Does that prove the deals caused the change?

No. We measured a correlation across seven sites and checked it against public reporting on those agreements. We cannot read anyone's contract, so we cannot say whether crawler access is written into the deal or simply follows from the relationship. Seven out of seven is strong enough to report and not strong enough to call proof.

Which crawler is blocked most often?

CCBot, operated by Common Crawl, at 90% of blocking sites. Common Crawl is a non-profit that has archived the web since 2011, long predating the current AI industry. ClaudeBot and ByteDance's Bytespider follow at 80%. OpenAI's GPTBot is blocked by 52%, ChatGPT-User by 38% and OAI-SearchBot by 33%.

What is the difference between GPTBot and OAI-SearchBot?

Broadly, GPTBot gathers content for training, while OAI-SearchBot supports ChatGPT's search feature, which cites and links to sources. Blocking the first and allowing the second is a coherent position for most publishers: no training, but still be quoted and linked. Blocking both removes you from the answers as well.

Should I block AI crawlers on my own site?

It depends entirely on how you make money. If people pay to read your content, an assistant that answers in your place is a competitor and blocking is defensible. If people pay to use your product, an assistant reading your documentation is helping you sell — which is why not one developer-tool or SEO company in our sample blocked anything at all.

Can I run this check myself?

Yes, and it is free. Add /robots.txt to any domain and read it. Look for the AI crawler names, check whether each has its own group with Disallow: /, and note who is missing. Doing this for the ten biggest sites in your industry takes about twenty minutes and tells you a great deal about their commercial position.

Read the file before you copy it

Every case study in this series has looked at something a company built. This one is about something companies accidentally published.

Nobody set out to disclose their negotiating position in a plain text file at the root of their domain. But contracts have technical consequences, technical consequences get written into robots.txt, and robots.txt is public by design.

So the next time you see a recommended AI blocklist doing the rounds, remember what you are looking at. It is not best practice. It is the residue of somebody else's negotiation, and the parts they left out are the interesting bit.

Start with your own file — run a free audit and see what you are currently telling crawlers. Most people are surprised.

Checked on 22 August 2026 across 49 domains, parsing robots.txt by user-agent group rather than by keyword match, and counting only site-wide Disallow: / rules. Full category and per-crawler figures are in the FAQ above.

Licensing agreements were checked against public reporting, not against contracts, and no claim is made about their terms. Both robots files and commercial arrangements change often; re-run the check before relying on it.

Blog