Perplexity
Perplexity searches the web in real time for each question and returns a conversational answer with numbered citations linking to the sources. It indexes sites with PerplexityBot, which it says is not used to train foundation models, and fetches pages for users with Perplexity-User, which generally ignores robots.txt.
How it sources answers
Perplexity describes itself as an AI-powered search engine that “searches the web to deliver accessible, conversational answers backed by verifiable sources,” with content “sourced from the web in real-time.” Each response “includes citations and links to original sources.”[1]
That makes Perplexity the clearest example of retrieval-first design: the cited pages are what the answer was built from (see how answer engines work). Perplexity does not publish how it ranks or selects sources.
Perplexity’s crawlers
| User agent | Purpose | Documented user-agent string |
|---|---|---|
| PerplexityBot | Surfaces and links websites in Perplexity search results; “not used to crawl content for AI foundation models” | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot) |
| Perplexity-User | Visits pages to answer a user’s question; “generally ignores robots.txt rules” because the user requested it | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user) |
Perplexity recommends allowing PerplexityBot “to ensure your site appears in search results” and publishes IP address lists for both agents, with instructions for WAF allow-listing.[2] Vendor-documented
Robots.txt disputes
In the Tow Center’s 2025 study, Perplexity’s free version correctly identified excerpts from National Geographic articles even though the publisher disallowed Perplexity’s crawlers, one of several signs the researchers cited that chatbots may bypass robots preferences.[3] Because Perplexity-User is user-triggered and generally ignores robots.txt,[2] a robots.txt block does not guarantee your content never appears.
How to be cited by Perplexity
- Allow PerplexityBot in robots.txt and in your CDN or WAF. Vendor-documented
- Answer the question in the first lines of the page and keep facts explicit and dated (see content structure). Practitioner practice
- Keep pages fresh with visible dates, since answers are built from real-time search (see freshness). Practitioner practice
- Be present on sources Perplexity tends to cite in your niche. Check by running your target questions and noting the cited domains (see measurement). Practitioner practice
Accuracy
Perplexity had the lowest error rate of the eight tools the Tow Center tested, yet still answered 37% of queries incorrectly.[3] Researchers also showed that adversarial text on product pages could promote low-ranked products in conversational search, and that the attacks transferred to perplexity.ai.[4] Independent study
Measurement
Perplexity does not offer a publisher dashboard. Track referrals from perplexity.ai in analytics and run a fixed prompt set on a schedule (see AI traffic in GA4).
Frequently asked questions
Does Perplexity use my content for training?
Perplexity says PerplexityBot is not used to crawl content for AI foundation models.
Can I block Perplexity?
You can disallow PerplexityBot in robots.txt, which removes you from its search index. Perplexity-User fetches triggered by users generally ignore robots.txt, so block by IP if you need to stop those.
Why does Perplexity cite my competitor?
Perplexity does not publish its ranking method. Compare the cited page with yours: it may answer the question more directly, be fresher or be easier to retrieve.
See also
References
Pages accessed September 28, 2026 unless a date is given. See all sources and our editorial policy.
- ^ "What is Perplexity?". Perplexity Help Center.
- ^ "Perplexity Crawlers". Perplexity documentation.
- ^ "AI Search Has a Citation Problem (Jaźwińska and Chandrasekar)". Columbia Journalism Review, Tow Center for Digital Journalism. Published March 6, 2025.
- ^ "Ranking Manipulation for Conversational Search Engines (Pfrommer, Bai, Gautam, Sojoudi)". Proceedings of EMNLP 2024, ACL Anthology. Published November 2024.