# What blocking AI crawlers costs a creator's site

By Tai Nguyen, ExDarkMatter. Published 2026-10-04; updated 2026-10-04.
Canonical: https://exdarkmatter.com/notes/what-blocking-ai-crawlers-costs/

**Short answer:** Less than many creators fear, and not always what they expect. In our study, 476 of 528 creator sites (90.15%) blocked no AI crawler at all. Where a site does block one, the effect depends on the bot: blocking a search crawler can keep the site out of that engine's answers, while blocking Google-Extended does not change Google Search.

## How many creator sites block AI crawlers?

Very few. In The Dark Archive, our study of long-form YouTube channels, we scored 528 websites that creators own and read the robots.txt file of each one. That file is where a site tells each crawler, by name, whether it may fetch pages. We checked it for five widely used AI crawlers: GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot. Most of these collect pages for AI models; the search crawlers described further down were not part of the count.

Of the 528 sites, 476 blocked none of them. That is 90.15%. Blocking an AI crawler is the exception among creators, not the rule.

We cannot tell from robots.txt alone whether a block was a choice. Many sites inherit their rules from a website template, a hosting default or a security plugin, and the owner may never have looked at the file. So the small group that blocks should not be read as a group of creators who decided to opt out.

## What is the difference between a training crawler and a search crawler?

AI companies do not run one crawler each. They run several, for different jobs, and a site can allow one while blocking another. The distinction that matters most for a creator is between crawlers that collect pages to train models and crawlers that fetch pages to answer searches.

OpenAI documents OAI-SearchBot as the crawler that surfaces websites in ChatGPT's search features, separate from its other bots. Google treats Google-Extended as a control for its AI models that stands apart from the crawler behind Google Search. Anthropic documents Claude-SearchBot as the crawler that indexes sites for search in Claude.

A creator who blocks "AI" in one sweeping rule usually blocks both kinds at once. That may not be what they meant. Many creators who object to their work training a model would still like an assistant to cite them, with a link, when someone asks about a topic they cover. Those are two different decisions, and the vendors' documentation lets a site make them separately.

## What does each block actually change?

Here is what the three vendors say, in their own documentation, about blocking their search-related crawlers:

- **ChatGPT.** OpenAI states that a site opted out of OAI-SearchBot is not shown in ChatGPT search answers, although it can still appear as a navigational link. A site that wants to be cited in ChatGPT search needs that crawler allowed.
- **Google.** Google states that Google-Extended "does not impact a site's inclusion in Google Search", and that it is not a ranking signal. Blocking it changes how Google may use the pages for its AI models, not whether they appear in Search.
- **Claude.** Anthropic states that disabling Claude-SearchBot stops it indexing the site for search, which it says may reduce the site's visibility and accuracy in user search results.

The pattern is consistent. Blocking a search crawler costs visibility in that engine's answers. Blocking a crawler that only governs model use does not remove the site from search.

## Can a site that blocks AI crawlers still be cited?

Yes, depending on the engine. Because Google-Extended does not affect inclusion in Google Search, an answer grounded in Google Search can still cite a site that blocks it. The same site may be missing from ChatGPT search if it also blocks OAI-SearchBot. One site can be visible in one assistant and absent from another, purely because of a few lines in a text file.

The full report includes a case like this from our own sample, along with the results of our citation test.

## So what should a creator do?

Start by looking. Open your own robots.txt file, at your domain followed by /robots.txt, and read which crawlers it names. Then decide bot by bot rather than with one rule:

- **If you want to be cited in AI search answers,** keep the search crawlers allowed: OAI-SearchBot for ChatGPT, Claude-SearchBot for Claude, and Googlebot for Google Search.
- **If you object to model training on your pages,** the training controls are separate, and blocking them does not have to cost you search visibility on every engine.
- **If you never chose your rules,** find out where they came from. A template or a plugin may be making the decision for you.

Access, though, is rarely the real gap. Most creator sites already let every AI crawler in, and most of them still give those crawlers very little to read. The bigger problem is that so few creators keep a written version of their videos at all, which is the subject of [how many YouTube channels keep a written archive](https://exdarkmatter.com/notes/how-many-channels-keep-a-written-archive/). The machine readability score in the [glossary](https://exdarkmatter.com/glossary/) describes what a crawler-friendly page looks like.

## Where are the full numbers?

This note covers crawler access only. The per-crawler counts, how blocking varies by tier and the citation results are in [The Dark Archive report](https://exdarkmatter.com/report/).

To check how readable your own channel's site is, [run the free Dark Matter Check](https://exdarkmatter.com/check/).

## Sources

- [OpenAI: overview of OpenAI crawlers](https://developers.openai.com/api/docs/bots), retrieved 2026-09-30
- [Google Search Central: Google's common crawlers](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers), retrieved 2026-09-30
- [Anthropic: does Anthropic crawl data from the web, and how can site owners block the crawler](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler), retrieved 2026-09-30
