Executive summary
Indexability is binary and non-negotiable: a page not in the index cannot rank regardless of how good it is. This guide goes beyond reading the Google Search Console coverage report to a systematic reconciliation of intended-index versus actual-index state, a diagnosis playbook for each exclusion reason, and a discipline for deliberately excluding the pages that should never compete.
What this helps you decide
Which pages belong in the index and how to fix each specific reason a page you want indexed is being excluded.
Business problem
Pages that never make the index earn nothing, yet indexation problems are invisible in traffic reports and often go undiagnosed for months. Meanwhile low-value pages that should be excluded dilute site quality signals and waste the discovery that priority pages need.
Step-by-step process
-
1
Define the intended index set
Before diagnosing anything, decide which URL patterns should be indexable and which should not. Faceted filters, internal search results, thin tag archives and staging paths usually should be excluded (SS-DE-024); canonical product, category and content pages should be in. This intended set is the yardstick every diagnosis measures against.
-
2
Reconcile intended against actual
Pull the Search Console index coverage and page indexing reports and compare against your intended set. You are looking for two failures: valuable pages excluded (a leak) and low-value pages indexed (bloat). Both distort quality signals and demand different fixes.
-
3
Diagnose each exclusion reason precisely
Excluded-by-noindex, crawled-not-indexed, discovered-not-indexed, duplicate-without-canonical and soft-404 each have distinct causes. Group excluded valuable pages by their exact reason so you fix root causes rather than symptoms.
-
4
Resolve crawled-but-not-indexed as a quality signal
When Google crawls a page and declines to index it, treat it as a verdict on quality or redundancy, not a technical bug. Strengthen the content, consolidate near-duplicates, and improve internal links so the page earns its place rather than forcing it in.
-
5
Fix canonical and duplicate confusion
Where duplicate or parameterised URLs split indexation (SS-PT-14), set self-referential canonicals on the version you want indexed and canonicalise variants to it (SS-DE-045). Ensure canonicals, internal links, sitemaps and redirects all point to the same chosen URL.
-
6
Deliberately exclude the bloat
Apply noindex or robots controls to the low-value patterns you identified, and confirm they are still crawlable long enough for the directive to be seen. Reducing indexed bloat concentrates the site's perceived quality on pages that can convert.
-
7
Keep sitemaps and signals consistent
Submit only canonical, indexable, 200-status URLs in XML sitemaps. A sitemap full of redirected, noindexed or excluded URLs sends mixed signals and slows the resolution of legitimate pages.
-
8
Monitor index state as a standing metric
Track indexed-valid count against your intended set over time. A sudden divergence is an early warning of a template regression, an accidental sitewide noindex, or a canonical misconfiguration long before traffic reflects it.
Worked example
Checklist
- Intended index set defined by URL pattern before any diagnosis
- Actual index state reconciled against intended, listing both leaks and bloat
- Every excluded valuable page grouped by its precise exclusion reason
- Crawled-not-indexed pages addressed as quality or duplication issues, not forced in
- Canonicals self-reference the chosen URL and all signals agree
- Low-value patterns deliberately noindexed and confirmed crawlable to see the directive
- XML sitemaps contain only canonical, indexable, 200-status URLs
Common mistakes
- Treating crawled-but-not-indexed as a technical bug to override rather than a quality verdict to answer
- Blocking a page in robots.txt while also applying noindex, so the noindex is never crawled and the page lingers indexed
- Submitting redirected, noindexed or duplicate URLs in the sitemap and muddying the signals you need Google to trust
- Never defining an intended index set, so you cannot tell a legitimate exclusion from a damaging leak
Best practices
- Fix indexation before investing in on-page or link work, because a page outside the index returns zero on every other optimisation
- Make canonical, internal-link, sitemap and redirect targets agree on one URL per piece of content so signals never contradict
- Treat rising indexed bloat as a quality liability and prune deliberately rather than letting archives and parameters accumulate
- Watch indexed-valid count as a leading indicator; a sudden drop usually signals a template or directive regression before traffic moves
- Never robots-block a URL you want deindexed until after the noindex directive has been crawled and honoured
Troubleshooting
30-minute experiment
KPIs to track
- Indexed-valid page count against the intended index set
- Number of valuable pages recovered from excluded states per month
- Ratio of indexed bloat pages trending down after deliberate exclusion
FAQs
Should I fix indexing before working on content or links?
Yes. A page that is not indexed cannot rank no matter how strong its content or backlinks, so indexation is the prerequisite that every other optimisation depends on.
Is more indexed pages always better?
No. Indexing thin, duplicate or low-value pages dilutes the site's perceived quality and wastes discovery. The goal is the right pages indexed, not the most pages indexed.
Why is Google ignoring my noindex tag?
Almost always because the URL is also disallowed in robots.txt, so the crawler never reaches the page to read the noindex. Allow crawling until the page is deindexed, then apply crawl-level controls.
Recommended next steps
Where this fits - and what's next
The SearchScore path from a problem you feel to visibility you can measure.