6 min read

One AI engine shows its sources ten times more than the rest

The short version

We collected every source URL six AI engines handed back across 110 scans: 10,201 URLs in total. 73.8% of them came from Perplexity alone. Claude and DeepSeek returned 7.2% each, ChatGPT 4.2%, Grok 3.9%, Gemini 3.8%. The gap is not that Perplexity reads more of the web. It is that Perplexity tells you what it read, and the others largely do not.

Every AI engine that answers a question about your business has read something to answer it. Which pages it read is the single most actionable fact in AI visibility: it is the difference between "improve your content" and "get cited on these eleven pages". So we measured how much each engine is willing to say.

10,201
source URLs collected
73.8%
came from Perplexity
3.8%
the quietest engine

What each engine hands back

One row per engine, across the same questions, the same accounts and the same scans. "Hosts" counts distinct domains after unwrapping any redirects, so it reflects real publishers rather than URL formatting.

EngineSource URLsShare of all URLsDistinct hostsURLs per host
Perplexity7,53073.8%1,9973.8
Claude7307.2%3082.4
DeepSeek7307.2%3142.3
ChatGPT4334.2%1822.4
Grok3933.9%1682.3
Gemini3853.8%2041.9

Perplexity returns roughly ten times as many source URLs as the quietest engine, and covers six times as many distinct publishers. On a single answer it will routinely name a dozen pages; the others name two or three, or none.

The mistake this data invites

The obvious reading is that Perplexity researches harder. That does not follow, and we want to be first to say so.

What we can observe is what an engine returns in its response. An engine that consults ten pages and cites one is indistinguishable, from the outside, from an engine that consulted one. Every number above measures disclosure, not effort, and nothing in this dataset can separate the two. Anyone telling you these figures show which AI "does more research" is reading something into them that is not there.

That matters commercially, because the actionable half is the same either way: if an engine will not tell you what it read, you cannot go and get cited on it.

The Gemini number that was wrong, and how

Until this week our own source map showed Gemini contributing one domain. That was our bug, not Gemini's behaviour, and it is worth explaining because anyone else measuring this will hit it.

Gemini does not cite pages directly. It cites through vertexaisearch.cloud.google.com/grounding-api-redirect/<token>, a wrapper that forwards to the real page. Count hostnames naively and every Gemini citation collapses into a single Google domain. Follow the redirects and the real figure appears: 204 distinct publishers, comparable to ChatGPT's 182 and Grok's 168.

There is a deadline attached. Those tokens expire: in our data every one resolved under 30 days old, and every one over 40 days was dead, returning 404 with no way to recover the destination. If you are collecting Gemini citations and not resolving them within about a month, you are keeping URLs that will never again tell you anything.

How this was measured

The figures move as the corpus grows. This is the reading on 12 August 2026, and we publish the method so it can be checked rather than taken on trust. Every SearchScore customer can export the same cell-level data for their own account, with the source URLs and the retrieval mode attached, and recompute all of it.

What to do with it

If you are trying to appear in AI answers, the practical consequence is that Perplexity is where the map is. It will tell you which pages win a question, which makes it the cheapest place to learn what an entire category's engines are rewarding. The others may well be reading the same publishers; they simply are not saying.

See which sources the engines use for your own category →