How Perplexity discovers and retrieves websites

This page covers the discovery layer: how PerplexityBot finds pages, how robots.txt and server access affect eligibility, and what happens before source selection. For the separate question of why one retrieved page earns a footnote and another does not, use our citation-selection guide.

Video transcript

Perplexity works nothing like ChatGPT, and most people optimise for the wrong one. Every Perplexity answer is a live web search that shows its sources. That changes everything. First, it's live. Perplexity retrieves fresh pages on every single query. Second, sources are visible. It shows exactly who it cited, no hiding. Third, freshness wins. Recent, relevant pages get pulled first. Save this if you want Perplexity to cite you. Full guide free on SearchScore.

Key takeaway: PerplexityBot is Perplexity’s documented search crawler. Keeping it able to reach public pages preserves the discovery path, but discovery is only the first stage; which retrieved sources Perplexity cites is a separate selection problem.

Perplexity numbers every source it uses, which makes it the easiest engine to audit. Our walkthrough on finding the sources AI cites uses it as the starting point for building a citation target list.

What makes Perplexity different

Unlike ChatGPT - which is primarily a conversational AI that sometimes browses the web - Perplexity is built from the ground up as a search engine. Every query on Perplexity triggers a live web retrieval, and every answer comes with cited sources displayed prominently.

This makes Perplexity a high-priority target for GEO work. If someone is using Perplexity, they are almost always going to see sources - and if your website is not one of them, you are missing visible attribution.

How Perplexity retrieves content

Perplexity uses PerplexityBot as its documented search crawler. Check that it is not accidentally blocked if you want public pages eligible for Perplexity search. OAI-SearchBot belongs to OpenAI and should not be listed as a Perplexity crawler.

For each query, Perplexity:

1. Searches the web using its own index (built from PerplexityBot crawls)

2. Retrieves the top candidate pages in real time

3. Reads and extracts the most relevant content from those pages

4. Synthesises an answer and displays 5 to 6 source cards

What happens after discovery

Once a page is discoverable, Perplexity still has to decide whether it is relevant enough to retrieve for a specific question and whether any passage is useful enough to cite. That later selection stage involves different considerations from crawler access. Rather than duplicating them here, see How Perplexity decides which sources to cite for relevance, freshness, answer structure and source-selection analysis.

Perplexity discovery checklist

- Confirm PerplexityBot is not accidentally blocked in robots.txt; review other providers’ crawlers separately

- Make important public pages return a normal 200 response to PerplexityBot

- Check CDN/WAF rules as well as robots.txt

- Keep XML sitemaps and internal links clean so important URLs are discoverable

- Make sure primary content is present in the fetched HTML rather than hidden behind an inaccessible interaction

- Then evaluate source selection separately using the citation guide

Key insight: Treat discovery and citation as two different checks. A page can be crawlable yet never retrieved for the target question, or retrieved yet lose the footnote to a clearer source. Separating those stages makes the diagnosis much more useful.

Back to pillar

- What is GEO? The Complete Guide →

S

Ronnie Huss

GEO Research & Analysis

The SearchScore editorial team researches and writes about generative engine optimisation, AI search visibility and the signals that determine whether your website gets cited by ChatGPT, Perplexity and Google AI Overviews.

Sources & Further Reading

- Perplexity – PerplexityBot crawler documentation

- Academic research – GEO: Generative Engine Optimization (Aggarwal et al., arXiv)

- SearchScore SAVI Report, Q3 2026 (Vol. 3, 1,000,000+ sites analysed)

Check your AI visibility

Enter your URL at SearchScore for a free AI visibility Score. See how ChatGPT, Perplexity and Google AI see your site - and exactly what to fix.

Frequently asked questions

What crawler does Perplexity use?

Perplexity documents PerplexityBot as its search crawler and says it follows explicit robots.txt limits. OAI-SearchBot is OpenAI's crawler, not a Perplexity crawler. If Perplexity search visibility matters, check PerplexityBot access and your wider indexability rather than allowing unrelated user agents.

How does Perplexity decide which sources to show?

Perplexity does not publish a complete ordered citation-ranking formula. Relevance, freshness for time-sensitive queries, clear factual passages and source quality are sensible optimisation targets, but their weight varies by query and should be measured rather than treated as fixed rules.

Is Perplexity different from ChatGPT for source selection?

Yes. Perplexity is built primarily as a search tool and retrieves web content for almost all queries. ChatGPT only uses Browse for queries where current information is needed. Perplexity also tends to show more sources per answer and has a stronger emphasis on recency.

Part of Pillar Article - see all guides in this series →