19,998 Sites Block GPTBot While Allowing OpenAI Search
Across 347988 websites audited by SearchScore, OpenAI's model-training crawler GPTBot was blocked on 13.5%. Its search crawler OAI-SearchBot was blocked on 7.83%.
This is the cleanest comparison in the dataset because the two roles are documented by the same operator. OpenAI identifies GPTBot as a crawler for content that may be used to train its generative AI foundation models, while allowing OAI-SearchBot helps make content eligible to be discovered and cited in ChatGPT search.
GPTBot's block rate is 72.4% higher than OAI-SearchBot's in this sample. More importantly, the cross-tab shows the difference is often explicit within the same robots file: 19998 sites block GPTBot while allowing OAI-SearchBot, against 270 doing the reverse.
All nine measured tokens
The wider table is published as context, not as a binary training-versus-retrieval ranking. The roles are not uniform. For example, Google-Extended is a robots control token covering both Gemini training and grounding, not a separate HTTP crawler, while CCBot builds Common Crawl's general open-web dataset.
| Token | Conservative role label | Blocked | Rate |
|---|---|---|---|
| GPTBot | OpenAI model-training crawler | 46984 | 13.5% |
| CCBot | Common Crawl open-web dataset crawler | 46128 | 13.26% |
| Bytespider | ByteDance crawler; exact public purpose not relied on in the headline | 44595 | 12.82% |
| ClaudeBot | Anthropic model-development/training crawler | 44561 | 12.81% |
| Google-Extended | Google robots control token for Gemini training and grounding | 43827 | 12.59% |
| PerplexityBot | Perplexity search-indexing crawler | 28325 | 8.14% |
| OAI-SearchBot | OpenAI search crawler | 27256 | 7.83% |
| DeepSeekBot | Other measured crawler token; role not relied on in the headline | 24914 | 7.16% |
| GrqBot | Other measured crawler token; role not relied on in the headline | 24296 | 6.98% |
What this does not prove. This is a snapshot of the SearchScore audit corpus, not a random sample of the entire web and not a time series. Robots policy does not tell us why a rule was configured, whether the crawler obeyed it, whether content was used for training, or whether the site later appeared in an AI answer.
Methodology and role definitions
The study uses 347988 unique domains measured during the same 90 days ending 18 August 2026. Every measured token uses the same domain denominator. The headline is intentionally limited to GPTBot versus OAI-SearchBot because their different roles are explicitly documented by OpenAI.
For the wider table we use conservative labels. Anthropic says ClaudeBot collects public web content that could contribute to model training. Perplexity says PerplexityBot is for surfacing and linking sites in search results, not foundation-model training. Google says Google-Extended is a control token for both Gemini training and grounding. Common Crawl describes CCBot as the crawler that builds its open web archive. We do not rely on an exact Bytespider, DeepSeekBot or GrqBot role to support the headline.
- OpenAI publisher and search-crawler guidance
- Anthropic crawler guidance
- Perplexity crawler guidance
- Google crawler and Google-Extended guidance
- Common Crawl CCBot documentation
Correction record. An earlier SearchScore report stated that 7.8% of sites blocked GPTBot. Re-analysis over the larger 347,988-domain population found the correct GPTBot figure is 13.5%. The earlier 7.8% figure closely matches the measured OAI-SearchBot rate of 7.83%. SearchScore corrected the Q3 report on 18 August 2026.
Download and cite the data
The aggregate data behind the findings is available for reuse. Row-level domain data is not published.
For a sector or country-specific cut of this dataset, contact [email protected].