The question sounds simple, can Perplexity AI actually crawl your website, but the answer has real consequences for both your visibility and your control over your own content. The short version is yes, in more than one way, and what you can do about it is more limited than most people assume. Here is how Perplexity actually accesses pages, what robots.txt really controls, and what the whole thing means for your strategy.
The short answer
Perplexity can access your website through two distinct agents. PerplexityBot is the search crawler that builds and refreshes the index behind Perplexity’s answers, and it respects robots.txt. Perplexity-User is a user-triggered fetcher that visits a specific page in real time to answer a live question, and it generally ignores robots.txt because a human requested the fetch. So blocking the crawler limits whether you are indexed, but it does not necessarily stop every real-time visit. And there is a complication: Cloudflare has publicly reported undeclared, stealth crawling by Perplexity that bypassed blocks entirely. For most brands, though, the real goal is not to block Perplexity at all. It is to be crawlable and citable.
The two agents, and why the distinction matters
Per Perplexity’s own crawler documentation, the two agents do different jobs and follow different rules.
PerplexityBot is the automated search crawler. It reads and indexes pages so Perplexity can surface and cite them in answers. Perplexity recommends allowing it in robots.txt, and it follows those directives, which means if you disallow it, you reduce your chances of being indexed and therefore cited.
Perplexity-User is the real-time fetcher. When someone asks Perplexity a question and the system decides to read a specific page to answer accurately, Perplexity-User makes that visit. Perplexity states this agent generally ignores robots.txt, on the reasoning that a user directly requested the fetch, much like a person pasting your URL into the tool. So a robots.txt disallow stops the indexing crawler but does not reliably stop a user-triggered fetch.
| Agent | Purpose | robots.txt | What blocking it does |
|---|---|---|---|
| PerplexityBot | Builds and refreshes the search index | Respects it | Limits indexing, so fewer citations |
| Perplexity-User | Real-time fetch to answer a live query | Generally ignores it | Little effect, it acts on a user request |
That table is the crux. Many site owners add a robots.txt disallow expecting to be invisible to Perplexity, then are surprised to see their content still appear, because the real-time fetcher operates under different rules.
The Cloudflare complication
It gets thornier. Cloudflare published research alleging that Perplexity went further than its declared agents. According to Cloudflare’s report, Perplexity used not only its declared user agent but also an undeclared crawler impersonating Google Chrome on macOS, rotating through IP addresses outside its official range and switching autonomous system numbers to evade blocks. Cloudflare said that on test domains protected with both robots.txt disallows and firewall rules against Perplexity’s declared crawlers, querying Perplexity still returned detailed content from those restricted pages, behavior it measured across tens of thousands of domains and millions of requests per day. Perplexity has disputed the framing, but the takeaway for site owners is sobering: robots.txt is a polite request, not a lock, and stopping a determined fetcher usually takes firewall or bot-management rules rather than a text file.
So can you stop it?
Partly, and only with the right tools. A robots.txt disallow will limit PerplexityBot and therefore your indexing, which is the opposite of what most brands want. To actually restrict access, including the real-time fetcher and any undeclared crawling, you generally need server-level controls: a web application firewall, bot management, or a CDN that can identify and challenge these agents. Even then, as the Cloudflare findings show, a crawler that rotates IPs and impersonates a browser is hard to fully block. The honest position is that you have strong control over your indexing crawler and partial, effort-dependent control over everything else.
But should you block Perplexity at all?
For the large majority of brands, no. Blocking PerplexityBot removes you from the index behind Perplexity’s answers, which means you cannot be cited there, and AI answer engines are a growing source of discovery. Unless you have a specific reason, paywalled content, sensitive material, or a deliberate licensing stance, blocking is self-defeating. The strategic goal for most sites is the reverse of this whole question: make sure Perplexity can crawl you, and make your pages worth citing when it does.
What actually earns Perplexity citations
If the goal is visibility rather than blocking, the levers are clear and they are not about crawling mechanics. Perplexity favors fresh, well-structured, verifiable content and leans heavily on community discussion. A study of citation patterns found Reddit made up 46.7 percent of Perplexity’s top cited sources, reflecting how much it trusts the discussed, current web. And being named at all tracks with being talked about: Ahrefs studied 75,000 brands and found branded web mentions and YouTube mentions correlate far more strongly with AI visibility than backlinks. So once Perplexity can crawl you, the work is to answer specific questions clearly and recently, and to earn genuine mentions where it looks. The deeper question of whether Perplexity weighs authority is covered in are Perplexity answers based on domain rating, and the cross-engine playbook is in how to actually show up in Perplexity and ChatGPT search.
How to check whether Perplexity is accessing you
You can verify this yourself rather than guess. Two practical checks. First, look at your server logs or your CDN analytics for the user agents PerplexityBot and Perplexity-User, both of which identify with perplexity.ai in the user agent and can be confirmed by reverse DNS. Seeing PerplexityBot means the index crawler is reading you, seeing Perplexity-User means real-time fetches are happening on live questions. Second, ask Perplexity questions your pages should answer and watch whether it cites you, which tells you not just that it can access you but that it finds you worth citing. If you see crawler activity but no citations, access is not your problem, citability is. If you see neither, check that you are not accidentally blocking the crawler and that your pages are indexable.
What to do if you genuinely must restrict access
Some organizations do have legitimate reasons to limit AI access, sensitive material, paywalled content, or a deliberate licensing position. If that is you, understand that robots.txt alone is insufficient, because the real-time fetcher generally ignores it and undeclared crawling has been reported. Effective restriction requires server-level controls: a web application firewall, bot-management rules, or a CDN feature designed to identify and challenge AI agents. Even then, expect partial results against a crawler that rotates IPs and impersonates a browser, and revisit the rules periodically. The honest framing is that you can make access difficult and costly, but a determined fetcher is hard to fully exclude, so weigh the effort against the actual sensitivity of the content.
A note for publishers worried about content use
Publishers face a real tension: they want the visibility of being cited but worry about their content being used without traffic in return. Both concerns are valid. The pragmatic middle path for most is to stay crawlable for the index crawler so you remain eligible to be cited, while using server-level controls to limit the behavior you specifically object to, and to focus content strategy on being named and recognized even when the click does not come, since AI answers reduce clicks regardless. Treating Perplexity purely as a threat to block tends to cost more visibility than it saves, while treating it purely as free distribution ignores legitimate content-use concerns. Decide deliberately based on your business model rather than defaulting to either extreme.
Crawlability is table stakes, citability is the goal
Step back and the whole crawling question resolves into a simple priority. Being crawlable is table stakes, it is what makes you eligible to appear at all, much as Google says a page must be indexed and eligible for a snippet to show in its AI features. But eligibility is not the win. Plenty of crawlable pages are never cited because they do not answer a specific question clearly or are not discussed anywhere. So once you have confirmed Perplexity can read you, stop thinking about the crawler and start thinking about citability: clear answers, freshness, and earned mentions. The crawler question is where people start, and the citability question is where results actually come from.
A quick checklist
To put this into practice: confirm PerplexityBot is allowed in robots.txt if you want to be cited; check your logs for PerplexityBot and Perplexity-User to see real access; ask Perplexity your target questions to test whether you are cited, not just crawled; if you must restrict, use firewall or bot-management rules, not robots.txt alone; and if your goal is visibility, which it is for most, put your effort into clear, fresh, quotable answers and earned mentions rather than into blocking. That list keeps you focused on the outcome that matters, being cited, instead of on the mechanics of a crawler you mostly want to welcome.
What blocking actually costs you
Before you block anything, weigh what it costs. Removing yourself from Perplexity’s index means you cannot be cited in its answers, and AI assistants are a growing slice of how people discover brands and research purchases. For most businesses that is real lost visibility, traded away to prevent a content-use concern that server-level rules could address more surgically. There are legitimate reasons to restrict, sensitive or licensed content, but defaulting to block out of a general unease about AI usually costs more than it saves. The more productive default for the large majority is to stay open to the crawler and compete to be cited, then restrict only the specific things you genuinely object to with the right tools. Visibility you give up is hard to win back, so make the decision deliberately rather than reflexively. And remember the decision is not all-or-nothing. You can welcome the index crawler so you stay eligible to be cited, allow the citations that bring real visibility, and still apply firewall or bot-management rules to the specific content you genuinely need to protect. That balanced posture gets you most of the upside of AI discovery without surrendering all control, and it is the right default for the large majority of sites that want to be found more than they want to hide.
The takeaway
Yes, Perplexity can crawl your website, through an index crawler that respects robots.txt and a real-time fetcher that generally does not, with reported stealth crawling on top. You can limit the polite crawler easily and the rest only with firewall-level effort. But for almost everyone the right move is not to fight the crawler, it is to welcome it and be worth quoting, because the brands that win Perplexity are the ones it can read and wants to cite, not the ones hiding behind a robots.txt line that the fetcher ignores anyway.
Frequently asked questions
Can Perplexity AI crawl your website?
Yes. Perplexity uses two agents: PerplexityBot, a search crawler that builds its index and respects robots.txt, and Perplexity-User, a user-triggered fetcher that visits a page to answer a live question and generally ignores robots.txt because a human requested it. So your pages can be accessed both ways.
Can you block Perplexity from crawling your site?
Partly. You can disallow PerplexityBot in robots.txt to limit indexing, but Perplexity-User, the real-time fetcher, generally ignores robots.txt, and Cloudflare has reported undeclared stealth crawling that bypassed blocks. Truly stopping access usually requires firewall or bot-management rules, not robots.txt alone.
Should you block Perplexity at all?
Most brands should not. Blocking PerplexityBot removes you from the index behind Perplexity’s answers, which means you cannot be cited there. Unless you have a specific reason to restrict access, the goal is the opposite: be crawlable, eligible, and quotable so Perplexity can cite you.
How do I get Perplexity to cite my pages?
Be crawlable, answer the specific question clearly and recently, and earn presence where Perplexity looks. SQSEO fans one seed keyword into the question-level prompts that trigger AI answers for free, so you target the queries worth being cited for instead of guessing.