How to find the sources AI cites about your market
Every AI answer is built from a handful of retrieved pages. Those pages are the real battleground, and most of them are not yours. This guide shows you how to pull the actual source URLs out of each engine by hand, aggregate them into a ranked target list, and decide where to turn the volume up.
SearchScore data: 216 of the 1,000,000+ sites analysed score AI-Ready (GEO 80 or above). That is 0.022%, or about 1 in 4,600. Source: SearchScore methodology.
What actually decides whether AI mentions you?
An AI answer is not a ranked list of pages. It is a short piece of writing assembled from a small set of retrieved sources, usually somewhere between three and fifteen URLs. The model reads that set, then writes a summary of it.
This has a consequence most teams miss. Your visibility is not decided by your website alone. It is decided by the contents of that retrieved set. If the set contains six industry round-ups and a directory listing, and none of them name you, then you are invisible in that answer no matter how good your own site is. The model cannot recommend what it did not read.
So the useful question is not “how do I optimise my site for AI”. It is “which specific pages is the AI reading when someone asks about my category, and how do I get into them”.
That list is knowable. Every major engine will show you its sources.
How to extract the source list by hand
Do this manually at least once. It takes about twenty minutes and it will change how you think about the problem.
Step 1: write decision-stage questions
Ask what a buyer asks at the point of choosing, not what they ask when idly curious. “Best commercial roofing contractor in Leeds” is decision-stage. “What is commercial roofing” is not.
Do not put your own brand name in the question. A branded question inflates your apparent citation rate and tells you nothing, because you are handing the model the answer. Ask the question your prospect asks before they have heard of you.
Five to ten questions is enough to start.
Step 2: ask each engine with web access on
Run each question through the engines separately, with web search or grounding active:
- Perplexity is the most transparent. It numbers every source inline and lists them under the answer. Start here.
- ChatGPT displays clickable source links inline. Since the May 2026 links update it links out rather than only naming sites, which makes the source list far easier to read than it was.
- Google Gemini exposes grounding sources, though it sometimes routes them through a redirect rather than showing the destination directly.
- Claude, Grok and DeepSeek all surface references when their search tools are active.
The ChatGPT change is worth understanding, because it altered what being cited is worth. Before 7 May 2026, ChatGPT largely named sites without linking to them. After it began showing clickable links, a study by Botpresso of 52 websites found ChatGPT referral sessions rose from 183,497 in the six weeks before the change to 363,089 in the six weeks after, an aggregate increase of 97.9%. Homepages moved most sharply, from 7,990 sessions to 113,183, lifting their share of ChatGPT referral traffic from 4.4% to 31.2% (Botpresso, 2026).
Read the distribution before you get excited, though. In that same study 36 sites gained, 14 declined and 2 were flat, and the single biggest winner accounted for 64% of the net increase while the top five accounted for 91%. Being in the retrieved set is winner takes most. That is precisely why knowing which sources are in the set matters more than knowing your average position.
One important caveat. An engine only returns sources when it genuinely retrieved from the live web. If a model answers confidently with no citations at all, that is a finding in itself: it is answering from training memory, describing your market as it existed whenever that data was collected. You cannot influence that answer this quarter, only the next training cycle.
Step 3: record every URL, not just your own
This is the step people get wrong. The instinct is to check whether you were mentioned and stop there. That reduces a rich dataset to a yes or no.
Record every source URL the engine returned, for every question, along with which engine returned it. You want the full picture of what the models are reading about your category.
Step 4: aggregate by domain and count
Now turn the raw URLs into something you can act on. For each unique page, count:
- How many times it was cited across all your questions.
- How many different engines cited it. A page cited by four engines is part of the shared consensus about your market. A page cited once by one engine is noise.
- Whether you own it.
- Whether it names a competitor.
A page cited repeatedly, by multiple engines, that you do not own and that names a competitor instead of you, is the single most valuable row in your spreadsheet.
What to do with the list
Sort every source into one of three buckets. Each has a different move.
Bucket one: your own pages that are already cited
These are proven. The retrieval system can find them, chunk them and extract from them. That is not true of most of your site.
Do not rewrite them. Strengthen them: keep the data current, add the specifics a model can quote, expand the sections that are actually being pulled. Then look at what these pages have in common structurally, and apply that pattern to the pages you wish were cited. Your own citation record is the most reliable style guide you will ever get. See how to write for AI citation for the underlying structure.
Bucket two: third-party pages that already mention you
A directory entry, a round-up you scraped into years ago, a comparison post, a forum thread. You are in the retrieved set already, but often as a single line among twelve competitors.
This is where you turn the volume up. Claim the listing. Expand the entry from one line to a full profile. Refresh outdated details. Get the review count up. Ask for the ranked position to reflect reality. You are not trying to get cited here, you are already cited. You are trying to be described better when you are.
This is usually the fastest win available, and it is almost always ignored, because it does not feel like content work.
Bucket three: pages that cite competitors but not you
Your target list. These are pages the models already trust, already retrieve, and already use to answer questions in your category. They are simply describing a market you are not in yet.
Work them in citation-frequency order. Getting onto one page that four engines cite is worth more than ten guest posts nobody retrieves. This is what digital PR for GEO is actually for: not links, but presence in the retrieved set.
The competitor angle most teams miss
Run the exercise a second time, but track the sources cited in answers that recommend your closest competitor.
You now have the specific pages feeding the model’s preference for that rival. Not a guess about their strategy, the actual evidence trail. If the same trade publication and the same two directories show up behind every answer that names them, you know precisely what to go and earn.
That is a competitive intelligence asset, and it is sitting in plain sight in every answer.
Where the manual method stops working
The manual pass is genuinely useful and you should do it. It also has a hard ceiling, for three reasons.
Volume. Six engines across ten questions is sixty answers per pass. Each returns its own source list. Doing that properly by hand is most of a day.
Variance. AI answers are not deterministic. Ask the same question twice and you can get different sources. One manual pass is a snapshot, not a measurement, and you cannot tell a real change from run to run noise without repeat sampling.
Drift. Source sets change as models re-crawl and re-index. The round-up that carried your category in May may be replaced by a different one in August. A single audit tells you where you stood on one day.
Which means the question is not whether to track sources, but whether to do it once or continuously.
How SearchScore Tracker does this automatically
SearchScore Tracker runs the workflow above on a schedule. It asks your questions across all six engines, ChatGPT, Gemini, Claude, Perplexity, Grok and DeepSeek, and keeps the source data rather than discarding it.
What you get back:
- Top Sources. Every page the engines cited when your brand came up, ranked by citation frequency, with the number of engines that cited each one and a flag showing which pages are yours.
- Competitor sources. The same aggregation scoped to answers that recommend a named competitor, so you can see the pages feeding their recommendations.
- Change over time. Sources gained and lost between scans, so drift shows up as an event rather than a surprise.
Scans run weekly on Founders and Starter, and daily on Pro for those who opt in. Starter is £20 a month with a 14 day trial.
One point of honesty about the data: source URLs only exist for web-grounded answers. Where an engine replies from training memory, there is no source list to report, and Tracker shows that rather than inventing one. Knowing that a model is answering you from memory is itself actionable, because it tells you the fix is brand presence over time rather than a page you can publish this week.
Start with what you can see today
You do not need a tool to begin. Open Perplexity, ask the question your best customer asks before they know you exist, and read the sources.
If your competitors are in that list and you are not, you have found your next quarter of work, and you have found it in about ninety seconds.
To see where your own site stands on the 250+ signals that decide whether you are retrievable in the first place, run a free SearchScore audit. It takes 60 seconds and needs no email.
Frequently asked questions
How do I find out what sources ChatGPT uses about my business?
Ask ChatGPT a decision-stage question in your category with web search active, then open the citation links it displays inline. Since the May 2026 links update, ChatGPT shows clickable source links rather than only naming sites, so you can read the full source list directly. Record every URL, not just the ones that mention you. The sources that describe your market without naming you are the highest-value targets, because they are already trusted by the model and are currently working in a competitor's favour.
Which AI engines show their sources?
Perplexity is the most transparent and numbers every source inline. ChatGPT, Google Gemini, Claude, Grok and DeepSeek all expose sources when web search or grounding is active. The critical caveat is that an engine only returns source URLs when it actually retrieved from the live web. If the model answers from training memory alone, there are no sources to read, and a confident answer with no citations means the model is repeating what it absorbed months ago rather than reading your site today.
What should I do with the list of sources AI cites?
Sort the URLs into three groups. Pages you own that are already cited should be strengthened and kept current, because they are proven retrievable. Third-party pages that already mention you are where you turn the volume up, by deepening the listing, refreshing the data or expanding your entry. Third-party pages that cite competitors but not you are your target list for digital PR and outreach. That third group is the one most teams never see, and it is usually where the fastest gains are.
Can I do AI citation source tracking manually?
Yes, and you should do it once by hand before automating anything, because it teaches you what the data actually looks like. Manual tracking is one engine and one question at a time, which is genuinely useful for a spot check. It stops scaling quickly: six engines across ten questions is sixty answers per pass, and because AI answers vary run to run, a single pass is a snapshot rather than a trend. That is the point at which automated tracking earns its place.
Continue Reading
- AI visibility drift: why sites lose AI search rankings without changing anything
- ChatGPT says wrong things about my business: how to fix incorrect AI information
- Google AI Overviews update history: what changed and what it means for your visibility
- Run a free AI visibility check across every major AI engine