SearchScore Research · AI crawler policy

19,998 Sites Block GPTBot While Allowing OpenAI Search

Across 347988 websites audited by SearchScore, OpenAI's model-training crawler GPTBot was blocked on 13.5%. Its search crawler OAI-SearchBot was blocked on 7.83%.

Measured 2026-08-18 · 90 days ending 18 August 2026 · SearchScore audit corpus · Aggregate data released under CC BY 4.0
19,998domains block GPTBot while explicitly allowing OAI-SearchBot.
270domains show the reverse configuration.
42.6%of GPTBot blockers explicitly allow OpenAI's search crawler.
74 to 1the directional split between GPTBot-block/OAI-allow and the reverse.

This is the cleanest comparison in the dataset because the two roles are documented by the same operator. OpenAI identifies GPTBot as a crawler for content that may be used to train its generative AI foundation models, while allowing OAI-SearchBot helps make content eligible to be discovered and cited in ChatGPT search.

GPTBot's block rate is 72.4% higher than OAI-SearchBot's in this sample. More importantly, the cross-tab shows the difference is often explicit within the same robots file: 19998 sites block GPTBot while allowing OAI-SearchBot, against 270 doing the reverse.

All nine measured tokens

The wider table is published as context, not as a binary training-versus-retrieval ranking. The roles are not uniform. For example, Google-Extended is a robots control token covering both Gemini training and grounding, not a separate HTTP crawler, while CCBot builds Common Crawl's general open-web dataset.

TokenConservative role labelBlockedRate
GPTBotOpenAI model-training crawler4698413.5%
CCBotCommon Crawl open-web dataset crawler4612813.26%
BytespiderByteDance crawler; exact public purpose not relied on in the headline4459512.82%
ClaudeBotAnthropic model-development/training crawler4456112.81%
Google-ExtendedGoogle robots control token for Gemini training and grounding4382712.59%
PerplexityBotPerplexity search-indexing crawler283258.14%
OAI-SearchBotOpenAI search crawler272567.83%
DeepSeekBotOther measured crawler token; role not relied on in the headline249147.16%
GrqBotOther measured crawler token; role not relied on in the headline242966.98%

What this does not prove. This is a snapshot of the SearchScore audit corpus, not a random sample of the entire web and not a time series. Robots policy does not tell us why a rule was configured, whether the crawler obeyed it, whether content was used for training, or whether the site later appeared in an AI answer.

Methodology and role definitions

The study uses 347988 unique domains measured during the same 90 days ending 18 August 2026. Every measured token uses the same domain denominator. The headline is intentionally limited to GPTBot versus OAI-SearchBot because their different roles are explicitly documented by OpenAI.

For the wider table we use conservative labels. Anthropic says ClaudeBot collects public web content that could contribute to model training. Perplexity says PerplexityBot is for surfacing and linking sites in search results, not foundation-model training. Google says Google-Extended is a control token for both Gemini training and grounding. Common Crawl describes CCBot as the crawler that builds its open web archive. We do not rely on an exact Bytespider, DeepSeekBot or GrqBot role to support the headline.

Correction record. An earlier SearchScore report stated that 7.8% of sites blocked GPTBot. Re-analysis over the larger 347,988-domain population found the correct GPTBot figure is 13.5%. The earlier 7.8% figure closely matches the measured OAI-SearchBot rate of 7.83%. SearchScore corrected the Q3 report on 18 August 2026.

Download and cite the data

The aggregate data behind the findings is available for reuse. Row-level domain data is not published.

Journalistic citationSearchScore, AI Crawler Policy Study 2026, analysis of 347,988 websites, measured 18 August 2026. https://searchscore.io/research/ai-crawler-policy-2026/
Academic-style citationHuss, R. (2026). AI Crawler Policy Study 2026. SearchScore. https://searchscore.io/research/ai-crawler-policy-2026/

For a sector or country-specific cut of this dataset, contact [email protected].