How chatgpt decides which websites to cite

When ChatGPT search uses the web, it can cite sources from the retrieved results. OpenAI documents the discovery crawler, but it does not publish a complete ordered ranking-factor list for citation selection. This guide separates what is documented from the practical signals worth testing.

Video transcript

This is exactly how ChatGPT decides who to recommend. It's not random. Three things decide who makes the cut. First, it has to reach you. No crawl access, no citation. Second, direct answers win. Pages that answer the exact question get picked. Third, credibility decides. Trust signals settle the final shortlist. Learn how to get cited, free on SearchScore.

Key takeaway: OAI-SearchBot access is a documented eligibility factor for ChatGPT search discovery. The rest of citation selection is not published as a five-factor hierarchy, so use relevance, answer clarity, sourcing, entity consistency and structured data as testable optimisation inputs rather than claiming a fixed order.

How ChatGPT retrieves web content

ChatGPT answers in two ways. It can generate a response from training memory alone, or it can retrieve live content from the web when a query benefits from current information. Only the second produces citations.

When ChatGPT searches, it draws on the index built by its search crawler, OAI-SearchBot. It reads the retrieved pages, extracts the most useful information, synthesises an answer, and cites the sources it used. This is a different user-agent from GPTBot, which OpenAI uses for model training only: blocking GPTBot does not remove you from ChatGPT search, and blocking OAI-SearchBot does.

The selection of which pages to retrieve and which to cite involves multiple layers of evaluation.

The five factors ChatGPT uses to select citations

1. Accessibility - can OAI-SearchBot actually read your page?

The first filter is purely technical. If OAI-SearchBot is blocked in your robots.txt file, ChatGPT cannot cite your content at all. SearchScore’s interactive audit data (6,944 websites, July 2026) shows 6.9% block at least one major AI crawler - often by accident, through legacy User-agent: * rules that predate AI search.

Check your robots.txt: Visit yoursite.com/robots.txt and look for any rule that blocks OAI-SearchBot or uses User-agent: * with Disallow: /. Either will prevent ChatGPT from citing your site.

2. Relevance - does your page directly answer the query?

ChatGPT evaluates how well a page’s content matches the user’s specific question. Pages that directly address the query - with the answer stated clearly near the top - perform significantly better than pages that cover the topic tangentially.

This is why heading structure matters so much for GEO. A page with an H2 that reads “How does X work?” followed immediately by a clear answer is far more citable than a page where the same information is buried in the fifth paragraph of a section titled “Overview.”

3. Content quality and structure

ChatGPT’s retrieval system favours content that is factually accurate, well-structured and easy to parse. Specifically:

- Clear definitions stated explicitly (“X is defined as…”)

- Factual claims supported by data or cited sources

- Logical heading hierarchy (H1 > H2 > H3)

- Short, declarative sentences in key answer sections

- Lists and tables for structured information

4. Credibility signals

ChatGPT evaluates the trustworthiness of sources. Websites with strong brand authority - consistent presence across the web, mentions in reputable publications, verified entity information - are treated as more credible sources than anonymous or thin websites.

Organisation schema, named authors with credentials, and external brand mentions all contribute to the credibility profile that influences citation likelihood.

5. Structured data and optional agent files

Schema.org markup can provide machine-readable context about the business, author or page when a consumer parses it. Use FAQPage, Article, Organisation and other types only when they accurately describe visible content. OpenAI has not documented llms.txt as a ChatGPT search ranking or citation input, so treat it as optional agent-readiness housekeeping.

What ChatGPT does not use

Several traditional SEO factors that matter for Google rankings have little or no direct effect on ChatGPT citation:

- Raw backlink count - no published evidence that ChatGPT uses a simple backlink-count factor; links can still matter indirectly through the search and authority systems feeding retrieval

- Keyword density - semantic relevance matters, not keyword repetition

- Domain age - no published evidence that age alone is a direct citation factor

- Meta keywords - ignored entirely

This means a newer website with excellent GEO signals can outperform an established site that has never optimised for AI search.

A practical checklist to get cited by ChatGPT

- ☐ OAI-SearchBot is not blocked in robots.txt

- ☐ If you intentionally maintain llms.txt, it is valid, factual and current

- ☐ Organisation schema implemented

- ☐ Article schema on key content pages

- ☐ FAQPage schema on Q&A content

- ☐ Core answers stated in first 100 words of each page

- ☐ H2/H3 headings written as direct questions

- ☐ Brand mentioned consistently across third-party sources

Back to pillar

- What is GEO? The Complete Guide →

S

Ronnie Huss

GEO Research & Analysis

The SearchScore editorial team researches and writes about generative engine optimisation, AI search visibility and the signals that determine whether your website gets cited by ChatGPT, Perplexity and Google AI Overviews.

Sources & Further Reading

- OpenAI – OpenAI crawlers: OAI-SearchBot, GPTBot and ChatGPT-User

- Academic research – GEO: Generative Engine Optimization (Aggarwal et al., arXiv)

- SearchScore SAVI Report, Q3 2026 (Vol. 3, 1,000,000+ sites analysed)

Check your AI visibility

Enter your URL at SearchScore for a free AI visibility Score. See how ChatGPT, Perplexity and Google AI see your site - and exactly what to fix.

Frequently asked questions

Does ChatGPT always cite sources?

Not always. ChatGPT cites sources for queries where it retrieves live web content through its search layer. Answers drawn only from training memory, without a live retrieval step, carry no citations.

How does ChatGPT choose which pages to cite?

OpenAI documents OAI-SearchBot as the crawler relevant to ChatGPT search discovery. Beyond access, the exact weighting of relevance, content structure, authority and other factors is not publicly specified, so those should be treated as practical optimisation hypotheses rather than a fixed ranking formula.

How do I get ChatGPT to cite my website?

The key steps are: keep OAI-SearchBot access open if you want ChatGPT search discovery, make important pages indexable and clear, use accurate structured data where relevant, answer buyer questions directly, and build independent corroboration. llms.txt is optional and is not documented by OpenAI as a ranking input.

Part of Pillar Article - see all guides in this series →