How Claude decides which websites to cite

Claude surfaces a business in two different ways: recall from its training data, and live citations when its web search capability is switched on. Each path selects sources differently, and each needs different fixes. This guide explains exactly how both work, and what your pages need before Claude will quote them.

Key takeaway: Claude cites websites through two separate mechanisms. By default it answers from training data, where your brand’s historical footprint across the public web decides whether you are recalled. When web search is enabled (in claude.ai, the Anthropic API and Claude Code), Claude retrieves live pages and cites the ones it can reach, read and quote cleanly. You can influence the live path quickly; the training path is earned over time.

What are the two ways Claude cites a website?

Claude does not have a single index the way a search engine does. It can name your business in two quite different ways, and understanding the difference is the foundation of Claude visibility work.

Training data (the default). Claude answers primarily from what it learned up to its knowledge cutoff. If your brand appeared clearly and often across the public web before that date, Claude can recall and recommend you with no live lookup at all. There is no submit button and no way to edit this on demand: you cannot email Anthropic to add yourself. You earn a place in it over time through a distinct, well-referenced footprint.

Web search and tool use (the live path). Claude can search the web and retrieve pages for current answers. That makes current public sources actionable, but Anthropic publishes no universal “within weeks” citation window. Measure the same questions after material changes rather than treating live retrieval as a guaranteed fast lane.

A brand can win on one path and lose on the other. You might be recalled from training but never cited live because your current pages are unreadable to a crawler, or you might be absent from training yet pulled in live because your content answers the exact question being asked.

How does Claude’s training data decide which brands it recalls?

The training path rewards consistency and breadth. Claude learned about brands from the public web as it existed before its knowledge cutoff, so the question is not “how good is my website today?” but “how clearly and consistently did my brand appear across the web over time?”

The signals that shape training-data recall are the classic entity and authority signals:

None of this moves quickly. That is exactly why the live path matters more for most businesses working on Claude visibility right now.

How does Claude’s web search choose which pages to cite?

When Claude searches the web mid-conversation, the selection process looks much closer to the retrieval systems used by other AI engines. In practice, four filters decide whether your page makes it into a cited answer:

1. Can Anthropic’s crawlers reach the page?

Anthropic documents three agents that honour robots.txt: Claude-SearchBot for search, Claude-User for user-directed retrieval and ClaudeBot for model training. Blocking Claude-SearchBot or Claude-User can reduce live visibility; ClaudeBot is a separate training choice. SearchScore’s July 2026 sample found 6.9% of sites blocked at least one major AI-related crawler or control checked by the audit.

2. Is the content actually in the HTML?

Crawlers do not reliably execute JavaScript. If your key content only exists after client-side rendering, Claude may fetch your page and find nothing usable. Server-rendered content is a baseline requirement.

3. Does the page answer the question directly?

Claude picks the source it can lift a clean, self-contained answer from. A page that states the answer plainly, near the top, under a heading that matches the question, is far more citable than a page where the same information is spread across five paragraphs of build-up. Across 850,000+ sites in the SAVI Report (April 2026 edition), the average on-page structure score is just 23.1/100: most pages simply do not contain a passage an AI engine can quote cleanly.

4. Can Claude resolve who you are?

Entity clarity matters, but no single markup type guarantees attribution. Keep visible brand details consistent, use accurate Organisation and Person markup where appropriate, and support important claims with reliable external evidence.

Where does Claude actually surface businesses?

A large share of Claude usage never touches the claude.ai chat window. Claude is one of the most heavily used models among developers, reaching people through the Anthropic API, inside other companies’ apps, assistants and internal tools, and through Claude Code on engineers’ machines.

Many of those products run their own retrieval: they search the web or their own content, then hand the results to Claude to answer from. So your business can be surfaced by Claude inside a customer support bot or research tool you have never heard of. You cannot optimise for each app individually, and you do not need to. The same fundamentals decide all of them: reachable pages, liftable answers and a well-referenced entity.

What does Claude not use?

Several traditional SEO signals have little direct effect on whether Claude cites you:

A practical checklist for getting cited by Claude

To see which of these you are missing today, run the free Claude Visibility Checker. It tests reachability for each Anthropic crawler separately, scores your citable structure, and returns a ranked fix list in about 60 seconds.

Related articles

Sources & Further Reading

Frequently asked questions

Does Claude browse the web?

By default, no: Claude answers from its training data. But claude.ai, the Anthropic API and Claude Code all offer an optional web search and tool use capability. When it is enabled, Claude retrieves live pages and cites them with source links, which is why your current site's crawlability and structure matter.

How is Claude different from ChatGPT for citations?

Both can combine model knowledge with live web retrieval. Anthropic documents Claude-SearchBot for search, Claude-User for user-directed retrieval and ClaudeBot for training; OpenAI documents OAI-SearchBot separately from GPTBot. Avoid reducing either product to one external index unless the provider documents that architecture.

Can I pay or apply to be cited by Claude?

No. There is no submission process and no paid inclusion. Visibility is earned through crawler access, quotable structure and a consistent, well-referenced brand footprint, the same foundations measured in a GEO audit.

Part of Pillar Article - see all guides in this series →