We scored more than a million websites for AI search. The best one got 86.6.
Across 1,000,000+ websites, 216 score AI-Ready on our 0 to 100 scale, which starts at 80. That is 0.022%, about 1 in 4,600. In July we added a tier above it, Verified, at 90 and above. It is still empty. The highest score recorded anywhere is 86.6. The reason is not that sites lock AI out: only 7.8% block GPTBot. It is that 54.3% publish no structured data at all. The door is open and the room is empty.
When someone asks ChatGPT, Gemini or Perplexity to recommend a business, the assistant does not return ten links. It names two or three. Being one of them is a measurable property of a website, and we have now measured it on more than a million of them. This is the Q3 2026 reading of our State of AI Visibility Index.
216 of more than a million is 0.022%, roughly one site in 4,600. That number is bleak but not new: every edition of this index has found AI-Ready to be vanishingly rare. What is new this quarter is the number next to it.
The ceiling nobody has cleared
In July we added a band above AI-Ready. Verified starts at 90 and marks the point where a site has not merely satisfied the signals but left nothing for a model to guess at: who published this, what the business is, what it does, why it should be believed, all stated in a form a machine can read without inference. It is a strict subset of AI-Ready, so it moved no site down and changed no counts.
After more than a million sites, nothing has reached it. The single highest GEO score in the corpus is 86.6. The gap from 86.6 to 90 looks small. It is the distance between a site an AI can cite and a site an AI cannot misread, and nobody on the open web has closed it.
This matters more than the low average, and it is a different kind of fact. An average tells you where the middle of the web sits, which is mostly a story about the long tail. A ceiling tells you what is achievable with the way sites are built today. The answer is that current best practice, executed well enough to beat a million competitors, still stops short of the bar.
| Band | Score range | Share of the index panel |
|---|---|---|
| Verified | 90–100 | 0% |
| AI-Ready | 80–89 | 0.03% |
| Strong | 60–79 | 13.14% |
| Emerging | 40–59 | 52.67% |
| Low Visibility | 20–39 | 32.28% |
| Invisible | 0–19 | 1.89% |
Just over half the web sits in Emerging, the 40 to 59 band. A site there is legible to a machine but gives it no reason to pick that site over any other. A third sits below 40, where an AI can barely use the page at all.
Access is not the bottleneck
The most common explanation for poor AI visibility is the one the data rules out first. Businesses reach for robots.txt, either to block the crawlers or to check they have not accidentally blocked them. That is the wrong file.
| What we checked | Share of sites | What it means |
|---|---|---|
| Block GPTBot | 7.8% | Permission is close to universal |
| Publish no JSON-LD at all | 54.3% | Nothing machine-readable on arrival |
| No Organisation schema | 67.6% | The engine cannot say what the business is |
| No llms.txt | 72.4% | No canonical summary of the site |
Seven times as many websites are unreadable as are closed. The crawler arrives, is welcomed in, and finds prose written for a human skim-reader with no structure attached to it. It leaves with nothing it can safely quote. Auditing robots.txt is the cheapest possible reassurance, and it answers a question almost nobody fails.
Technically fine, structurally invisible
The web has spent fifteen years getting the engineering right, and it worked. Not one site in our 20,000-site signal sample is missing HTTPS. That is what a solved problem looks like: a standard adopted to saturation, invisible because it works.
Now set that against the categories that decide whether a model can identify you at all.
| Category | Weight | Average score /100 |
|---|---|---|
| Technical | 12% | 70.8 |
| AI Citability | 18% | 70.4 |
| Topical Authority | 8% | 53.1 |
| AI Platform Readiness | 12% | 44.0 |
| Platform Optimisation | 4% | 41.5 |
| E-E-A-T Content | 24% | 38.5 |
| Structured Data | 12% | 26.7 |
| Brand Authority | 10% | 23.7 |
Structured Data averages 26.7 and Brand Authority 23.7. Those are not weak scores in a healthy distribution. They are the scores of a discipline nobody has started.
Two things in that table are worth staring at. The first is that E-E-A-T Content carries the heaviest weight in the index at 24%, and averages 38.5. The thing that counts most is the thing the web is worst at among the things anyone has attempted at all. The second is what is not the top weight. AI Citability scores 70.4, but its components are largely permissions a site already grants. Doing well on the category named after citation does not make a site citable.
The reason for the gap is structural, not lazy. Technical SEO had two decades and a tight feedback loop: fix it, watch the ranking move. AI citability has neither. Nothing on your analytics dashboard tells you that ChatGPT considered your page and could not work out who wrote it.
The schema cliff
Read the machine-readable layer signal by signal and you can watch a standard being adopted, then falling off a cliff.
| Signal | Sites missing it | What the engine cannot do |
|---|---|---|
| HTTPS | 0.0% | Nothing. This one is finished. |
| XML sitemap | 28.5% | Discover the full set of pages |
| Any JSON-LD | 54.3% | Read anything about the page as data |
| Open Graph image | 57.0% | Render the source in a citation card |
| Organisation schema | 67.6% | Say what the business is and is called |
| llms.txt | 72.4% | Find the canonical summary of the site |
| Author bio | 83.9% | Establish who is qualified to say this |
| IndexNow | 86.1% | Learn that the page changed |
| Person schema | 91.3% | Attribute the claim to a named human |
| FAQ schema | 96.6% | Match your page to a direct question |
None of these are difficult. Organisation schema is about fifteen lines of JSON. An author bio is a paragraph and a link. FAQ schema, absent from 96.6% of the web and the single cheapest way to make a page answer a question, is the largest unclaimed advantage in the entire dataset. The reason they are missing is not cost or difficulty. It is that nobody has told most site owners these signals are now load-bearing.
Why we are not comparing this quarter's average to the last
An honest note about the method, because it changes what you should take from this piece.
Averaging a growing corpus each quarter does not measure the web. It measures who joined the pool. When a dataset expands into the long tail the average falls, and none of that fall is a real decline. Our own earlier editions were also produced on an older scoring engine that ran a different set of checks and scored roughly three points more generously at the top of the range.
Put those together and a quarter-on-quarter average comparison would produce a headline that is an artefact in both directions. So we have not published one. From this edition the index runs on a fixed panel of 102,873 domains, frozen in July 2026 and stratified by tier so weighted aggregates reconstruct the whole population. The same domains are re-audited every wave, on the same method and the same scorer version. What changes between waves is the web, not the sample.
Wave 1, dated 1 August 2026, is the baseline: weighted mean GEO 45.24. It has nothing to be compared against yet, which is exactly what a baseline is for. The figure to track across editions is the AI-Ready count, which is measured consistently: 216 of 1,000,000+.
What to actually do about it
In order of score moved against time taken. The first four are a working week.
- Add Organisation schema to your homepage. Fifteen lines of JSON-LD naming the business, its URL, its logo and its social profiles. Missing from 67.6% of sites. Without it an engine infers your identity from prose, and it frequently infers wrong.
- Put a named, qualified author on every substantive page. A byline, a two-line biography, a link to a profile, and Person schema behind it. 83.9% have no author bio. E-E-A-T Content is the heaviest weight in the index.
- Answer real questions in FAQ schema. Absent from 96.6% of the web. Take the questions customers actually ask, answer each in under sixty words, and mark them up.
- Publish an llms.txt. A plain-text summary of what the site is and which pages matter. 72.4% do not have one.
- Build external corroboration. Brand Authority is the weakest category at 23.7. Consistent name, address and phone details across directories, a maintained company page, citations on sites you do not own. This one takes months, and it is the reason the ceiling sits where it does.
Read the full report: the SAVI Q3 2026 edition carries the full tier distribution, all eight weighted categories and the fixed-panel methodology. Or browse every edition on the SAVI index. You can score your own site free against the same 250+ signals.
Frequently asked questions
How many websites are AI-Ready?
216 of the 1,000,000+ websites analysed by SearchScore score 80 or above on the AI visibility scale, which is 0.022%, or about 1 in 4,600. No site has reached the Verified tier at 90 or above. The highest score recorded anywhere in the corpus is 86.6.
Does blocking GPTBot hurt AI search visibility?
It would, but almost nobody does it. Only 7.8% of websites disallow GPTBot in robots.txt, and the picture is similar for the other major AI crawlers. Access is not the constraint. The binding problem is that 54.3% of sites publish no structured data at all, so the crawler is let in and finds nothing it can safely quote.
What is the single most common AI visibility mistake?
Publishing no machine-readable identity. 54.3% of sites emit no JSON-LD whatsoever and 67.6% declare no Organisation schema, so an AI engine has to infer who the business is from prose. Structured Data averages 26.7 out of 100 and Brand Authority 23.7, the two weakest of the eight categories, while technical foundations average 70.8.
Why is FAQ schema worth adding?
Because 96.6% of websites do not have it, which makes it the largest unclaimed advantage in the dataset. AI assistants answer questions, and FAQ schema is the cheapest way to state, in a form a machine can read, that your page answers a specific one. Answer each question in under sixty words and mark it up.
Is the average AI visibility score getting worse?
We do not know, and we will not claim to. Earlier editions were scored on a different engine and on a corpus that more than doubled between readings, so a quarter-on-quarter average comparison would measure our own method rather than the web. From Q3 2026 the index runs on a fixed panel of 102,873 domains re-audited each wave, with a baseline weighted mean of 45.24. Real movement can be reported from wave 2 onward.