What AI engines actually evaluate when deciding whether to cite your site
ChatGPT, Claude, Gemini, Perplexity, Grok and DeepSeek do not publish one shared citation algorithm. The useful approach is to diagnose observable layers: can the relevant system discover the material, is it relevant and clear, is the claim supported, is the entity unambiguous, and are stronger sources winning the question?
SearchScore data: In the July 2026 interactive audit sample of 6,944 websites, 6.9% blocked at least one major AI-related crawler or control checked by the audit. That is a website-readiness statistic, not evidence of how often an engine cites a page. Source: SearchScore audit data, July 2026.
There is no universal AI citation scorecard
Different products use different retrieval paths.
- Google AI Overviews use Google Search systems and normal Search eligibility.
- ChatGPT Search uses OpenAI’s search stack, with OAI-SearchBot documented for search discovery.
- Claude documents Claude-SearchBot for search and Claude-User for user-directed retrieval.
- Perplexity operates its own search infrastructure.
- Grok and DeepSeek disclose less of their full source-selection architecture.
So SearchScore should not state that every engine applies the same six factors in the same order or with known weights.
What we can do is use a common diagnostic framework for observable problems.
1. Discovery and retrieval eligibility
Ask first whether the source can enter the relevant retrieval path.
That can mean:
- normal Googlebot indexability for Google Search and AI Overviews;
- OAI-SearchBot access for ChatGPT search discovery;
- Claude-SearchBot or Claude-User access for Claude live retrieval;
- PerplexityBot access where relevant;
- ordinary HTTP access, rendering and edge/WAF behaviour.
Training crawlers such as GPTBot and ClaudeBot are separate policy choices and should not be presented as live-search requirements.
A real access block is one of the clearest problems you can fix because it affects eligibility for that path. But the absence of a block does not mean a citation will follow.
2. Query relevance and coverage
A source has to answer the question the system is trying to solve.
For buyer queries, that may require more than a service page. The engine may be looking for:
- comparisons;
- evidence of price or availability;
- reviews and reputation;
- location or eligibility details;
- independent rankings;
- recent facts;
- a clear answer to a narrow sub-question.
This is why live source analysis matters. If the same directory, review site or publication keeps appearing, the next action may sit outside your own domain.
3. Factual clarity
Important information should be easy to locate, specific and self-contained enough to make sense in context.
Useful practices include:
- descriptive headings;
- direct factual statements;
- one clear idea per paragraph where practical;
- visible dates where freshness matters;
- explicit definitions and comparisons;
- avoiding vague marketing copy in place of the actual answer.
Question-shaped headings and answer-first paragraphs can be good editorial choices, but no major provider publishes a rule that they are mandatory citation syntax.
4. Evidence and source quality
Claims are easier to trust when they are supportable.
Depending on the topic, useful evidence can include:
- primary data;
- named authorship and relevant expertise;
- references to authoritative sources;
- transparent methodology;
- customer reviews;
- third-party coverage;
- official records.
Google’s people-first and quality guidance is relevant to Google Search surfaces. Other providers do not publish the same E-E-A-T framework, so do not present Google’s terminology as a universal LLM ranking system.
5. Entity consistency
A system can only represent a business accurately if the underlying facts are reasonably consistent.
Check:
- business name and trading name;
- website and canonical URL;
- address and service area;
- product/service descriptions;
- named people and roles;
- Organisation, Person, Product or LocalBusiness markup where appropriate;
- trusted external profiles and directories.
Structured data can reduce ambiguity for compatible consumers. It does not independently verify the claim and is not a universal citation requirement.
6. Freshness where the query needs it
Freshness matters when the answer itself is time-sensitive: prices, laws, product versions, opening hours, schedules, current leadership or recent research.
Update the page when the facts change and use honest publication/modification dates. Do not refresh timestamps purely to look newer, and do not assume a 2026 page automatically outranks an older authoritative source.
The most useful evidence: who is being used instead
A readiness audit can tell you that your site has an access, structured-data, content or entity problem.
It cannot see a provider’s private citation threshold.
Google’s AI Contribution pilot is a useful example of that measurement limit: third parties can observe surfaced sources and outputs, but Google has not published the formula it uses to value a source’s contribution during generation.
The fastest way to move from theory to evidence is therefore:
- run the same buyer questions;
- record who is named;
- record the source URLs/domains shown;
- compare those pages with yours;
- classify the gap;
- fix that specific layer;
- rerun the same questions.
That is the role of SearchScore’s Tracker: it measures the live output. The free audit remains the separate website-readiness diagnosis.
What not to call a ranking factor
Avoid presenting any of these as a universal direct AI citation lever:
- llms.txt;
- FAQPage markup;
- Organisation schema;
- question headings;
- a fixed word or chunk size;
- one exact author format;
- a fixed number of backlinks;
- a fixed publication cadence;
- a SearchScore score band.
All can be useful in the right context, but none is a provider-published universal citation threshold.
A practical order of operations
Use evidence to set the order rather than a canned factor hierarchy.
If there is a genuine discovery blocker: fix it.
If the cited sources contain information you do not: improve the underlying content or data.
If third-party sources dominate the answer: pursue legitimate visibility on those sources.
If the engine misidentifies the business: correct inconsistent public facts and markup.
If the site is technically healthy but still absent: do not keep polishing readiness signals indefinitely. Compare the live source set and solve the actual question-level gap.
The bottom line
AI source selection is not one hidden checklist that SearchScore can reverse-engineer into a universal formula.
The defensible model is diagnostic: eligibility, relevance, clarity, evidence, entity consistency and the competitive source set. Measure those, then verify the outcome in the engines themselves.
Sources & Further Reading
- Google Search Central - AI features and your website
- OpenAI - crawler and publisher guidance
- Anthropic - web crawler controls
- SearchScore - State of AI Visibility Index
Frequently asked questions
How does ChatGPT choose which sources to cite?
OpenAI does not publish a six-factor citation cascade. ChatGPT Search has documented discovery controls such as OAI-SearchBot, but source selection also depends on the query and the search/retrieval system. Diagnose access, relevance, factual clarity, evidence and the sources actually returned instead of assuming a fixed weighting.
Does domain authority matter for AI citation?
Authority can matter through upstream search systems, source reputation and independent corroboration, but no provider publishes a universal rule that authority is only a tiebreaker or that structure always beats it. Compare the sources returned for the exact question and identify the gap they satisfy.
What is the single most important factor for getting cited by AI?
There is no single universal factor. A hard discovery block can be decisive for one engine; a weak answer can matter for another query; a third-party source can dominate a recommendation question even when your own page is excellent. Fix the concrete bottleneck your audit and live-answer evidence reveal.
Continue Reading
- AI citations vs Google rankings: why ranking first on Google does not mean AI will cite you
- How AI knowledge graphs decide whether your brand exists (And what to do if you are missing)
- How AI retrieval works: what happens between your website and a ChatGPT answer
- Run a free AI visibility check on your site