Overview
Make sure ordinary search and the search/retrieval crawlers you depend on can discover and fetch important pages, while training choices remain separate.
Business problem
Important pages can be under-fetched because ordinary search crawlers or provider search crawlers are blocked, while teams often confuse search eligibility with separate training controls.
Decision supported
How to open access and steer crawl budget to what matters.
Inputs & outputs
Inputs
- Server logs
- robots.txt
- XML sitemaps
- Crawl stats
Outputs
- Crawler access matrix
- Crawl-budget steering plan
Step-by-step process
-
1
Map access
Map Googlebot/Bingbot plus provider search crawlers such as OAI-SearchBot and PerplexityBot separately from training crawlers such as GPTBot and control tokens such as Google-Extended.
-
2
Read logs
See where crawl budget actually goes.
-
3
Steer
Free budget from low-value URLs and point it at money pages.
-
4
Open selectively
Allow the search/retrieval crawlers whose discovery you want, and make training/grounding choices separately.
Maturity model
-
L1
Blind
No log or access visibility.
-
L2
Mapped
Crawler access documented.
-
L3
Steered
Crawl budget directed deliberately.
-
L4
Optimised
Access and budget tuned per crawler.
KPIs
- AI crawler access coverage
- Crawl budget on money pages
- Wasted-crawl %
Common mistakes
- Using a blanket 'AI crawler' policy without checking each token's documented role
- Letting infinite URL spaces burn ordinary crawl budget
- Sitemaps full of noindex URLs
- Blocking a search crawler when the intended goal was only to opt out of training
SearchScore insight
Recommended next steps
Where this fits - and what's next
The SearchScore path from a problem you feel to visibility you can measure.