Why your website isn't cited in Perplexity (and how to fix it)
If Perplexity rarely cites your site, separate discovery problems from source-selection problems. Check PerplexityBot access and WAF rules first, then page retrievability, freshness where the query demands it, answer clarity and external authority. Do not assume one blocker explains every absence.
Key finding: Perplexity is unusually search-centric and publishes substantial detail about its own web index and retrieval stack. SearchScore’s Q2 2026 SAVI benchmark found weak average page structure, and its July interactive sample found 6.9% blocked at least one major AI-related crawler or control checked by the audit. Those are readiness observations, not proof that they explain most missing Perplexity citations.
How Perplexity decides, in one paragraph
Perplexity describes its search stack as an internet-scale index with hybrid lexical/semantic retrieval, filtering and multi-stage ranking at document and passage level. That means a page can fail at discovery, retrieval or later ranking. The six checks below are a diagnostic sequence, not a published six-factor ranking formula.
Reason 1: PerplexityBot is blocked in robots.txt
PerplexityBot crawls and indexes the web so your pages can be discovered and cited as sources. Block it and you opt out of the pool Perplexity draws citations from, no matter how good your content is. Most blocks are accidental: legacy User-agent: * disallow rules, security plugins and CDN “block AI scrapers” toggles all catch it silently.
The fix: open yourdomain.com/robots.txt, look for rules affecting PerplexityBot or blanket disallows, and correct accidental blocks if Perplexity discovery is desired. Also check the WAF against Perplexity’s published IP ranges. Perplexity says crawler-setting changes can take up to 24 hours to reflect.
Reason 2: Perplexity-User is blocked at the network edge
Perplexity operates a second agent, Perplexity-User, for user-triggered page visits. Perplexity says it is not a crawler or training agent and generally ignores robots.txt. If those requests fail, inspect WAF, firewall or access-control rules rather than trying to solve the problem with an Allow directive in robots.txt.
The fix: check robots.txt for this agent separately. A blanket rule often catches both agents at once; make sure your allow rules cover each by name.
Reason 3: your content only exists after JavaScript runs
Perplexity quotes what it can read in the HTML it fetches. Client-side rendered pages hand the crawler an empty shell: navigation, a spinner, and none of your copy. You cannot be footnoted for content that is not in the fetched document.
The fix: view your raw page source and confirm your main content is present. Server-render or pre-render the pages you most want cited.
Reason 4: your pages are stale
Perplexity’s published search architecture explicitly filters stale content and optimises its index for freshness as well as completeness. That makes recency relevant for time-sensitive queries, but it does not mean every evergreen page loses visibility simply because it is older.
The fix: genuinely update your time-sensitive and cornerstone pages, show a visible last-updated date, and carry datePublished and dateModified in Article schema. Re-stamping dates without substantive changes does not survive contact with an engine that reads the page.
Reason 5: your answer is buried
The answer is assembled from passages Perplexity can lift and attribute right now. A page that states the answer plainly gets cited; one that buries it after three paragraphs of build-up gets skipped for a source that can be quoted cleanly. This is the single most common content failure in the SAVI data: sites average 23.1/100 on the structure signals that make a passage liftable.
The fix: lead each key page with one direct, self-contained sentence that answers the target question. Phrase headings as the natural-language questions people actually ask Perplexity. Use lists, tables and short paragraphs so the precise passage is easy to isolate.
Reason 6: weak authority on the topic
From the candidate pages, Perplexity leans toward reputable sources: domains that trusted sites mention and link to for the topic. If rivals have the third-party footprint and you do not, they win the footnote even when your content is comparable.
The fix: earn mentions and links from high-authority domains in your field, keep your entity data (Organisation and Person schema, consistent naming) clean so citations attach to the right brand, and build genuine topical depth rather than isolated pages.
The priority order for fixes
| Priority | Fix | Time | Why |
|---|---|---|---|
| 1 | Allow PerplexityBot | 10 minutes | No crawl, no source pool |
| 2 | Verify Perplexity-User at the WAF/edge | 10 minutes | User-triggered fetches generally ignore robots.txt |
| 3 | Server-render key content | Hours to days | Unfetchable content cannot be quoted |
| 4 | Refresh genuinely stale, time-sensitive pages | Varies | Perplexity filters stale results; relevance of freshness depends on the query |
| 5 | Rewrite answers to be liftable | 1-2 days | Quotability wins the footnote |
| 6 | Build topic authority | Ongoing | Decides ties against rivals |
How to confirm which blockers apply to you
The free Perplexity Visibility Checker checks PerplexityBot discovery, edge access, renderability, freshness signals, answer structure and entity clarity, then returns a prioritised fix list. Re-run the same buyer questions after fixes to see whether the real source set changes.
How to verify each fix landed
Perplexity’s query-time design makes verification unusually concrete. After each fix batch:
- Re-run the structured audit. Confirm the specific signal you fixed has moved: both agents green, rendering clean, dates detected, structure score up.
- Re-ask your test queries. Use the same five to ten buyer-realistic questions you baselined with, in fresh sessions, and record whether your domain enters the Sources list and where it sits.
- Watch the chosen page. When Perplexity starts citing you, note which page it picks and which sentence it quotes. That is direct feedback on what the engine considers your most liftable content, and a template for restructuring the rest.
Perplexity’s crawler settings can reflect within up to 24 hours according to its documentation, but recrawling, indexing and ranking have no universal public timetable. Use repeated measurements to establish when the change actually reaches answers.
Related articles
- How Perplexity decides which sources to cite →
- How to check your Perplexity visibility →
- Perplexity SEO guide: the full playbook →
- AI search for businesses: the full guide →
Sources & Further Reading
- Perplexity – PerplexityBot crawler documentation
- SearchScore SAVI Report, Q3 2026 (Vol. 3, 1,000,000+ sites analysed)
- Schema.org – Organisation structured data reference
- Academic research – GEO: Generative Engine Optimization (Aggarwal et al., arXiv)
Frequently asked questions
I rank well on Google. Why am I invisible on Perplexity?
Because Perplexity is not reading Google's index. It builds answers from its own crawl plus real-time search, so Google rankings and backlinks do not carry over. If PerplexityBot has never crawled you, or your page is stale, or the answer is buried, you can be page one on Google and absent from every Perplexity answer. The mechanics are covered in how Perplexity decides which sources to cite.
How quickly can I get back into Perplexity's answers after fixing these?
There is no guaranteed fastest-engine timetable. Perplexity says crawler-setting changes may take up to 24 hours to reflect, while indexing and ranking changes have no fixed public SLA. Re-run the same questions after recrawling rather than promising days or weeks.
Does Perplexity ever cite a site without crawling it first?
Perplexity-User supports user-triggered page visits. Perplexity says it is not a web crawler and generally ignores robots.txt because the user requested the fetch. If you intentionally want to block it, that is a WAF/access-control decision rather than a robots.txt optimisation.
Should I just block Perplexity to protect my content instead?
It is a legitimate policy choice. Blocking PerplexityBot in robots.txt opts the site out of Perplexity's crawler-based search discovery. Perplexity-User is a separate user-triggered fetcher and generally ignores robots.txt, so controlling it requires edge or access controls. Decide those two paths separately.
Does Perplexity cite social profiles and directories instead of websites?
It cites whatever page carries the most quotable, trusted answer, and for businesses with weak sites that is often a directory listing or review platform rather than the brand's own domain. If that is happening to you, it is reason 5 and 6 in this list: your own pages are not the most liftable source about you. Fix the structure and the citations follow the better source.