PerplexityBot and Perplexity-User: what is the difference?

PerplexityBot and Perplexity-User: what is the difference?

PerplexityBot and Perplexity-User: what is the difference?

THE SHORT ANSWER

PerplexityBot is the indexing crawler whose fetches feed the source pool Perplexity draws citations from, so allowing it is the precondition for appearing in a Perplexity answer. Perplexity-User is a live fetch triggered by a specific human request for a specific URL, closer in nature to a browser than a crawler. There has also been a documented public dispute about undeclared fetching, which is worth knowing before you rely on user-agent rules alone.

Perplexity is the assistant where citation is the interface rather than a footnote, which makes it the cleanest place to see whether your generative visibility work is landing. It is also the vendor whose crawler behaviour has attracted the most public scrutiny, so the mechanics deserve more care than a robots.txt block copied from someone else's blog post.

The two agents most people encounter serve opposite purposes, and the distinction is not cosmetic. Conflating them produces either a pointless block that costs you citations you wanted, or a false sense of control over fetching you never actually stopped. Both mistakes are common in robots.txt files written by an agency three years ago and never opened since.

The numbers, at a glance

  • PerplexityBot: the indexing crawler; if it cannot reach a page, that page cannot be cited in an answer

  • Perplexity-User: a live fetch on behalf of one person asking about one URL, which many operators classify as user traffic

  • Why it matters commercially: Perplexity shows sources prominently, so citation there converts to clicks at a higher rate than a buried reference

  • Open dispute: researchers and a major CDN publicly alleged fetching from undeclared agents; treat user-agent rules as necessary but not sufficient

The indexing crawler versus the on-demand fetch

PerplexityBot does what a search crawler has always done: discovers URLs, fetches them on its own schedule, and populates an index that gets queried later when somebody asks a question it can answer. Nothing about the timing is connected to any individual user. Block it and you are removed from the pool of things that can be retrieved, which is a straightforward and complete removal from the product.

Perplexity-User is the opposite shape. It exists because a person typed a URL or asked about a named page, and the system went to look. The request maps one-to-one onto a human intention, which is the argument vendors make for treating it as browsing rather than crawling. Whether that argument should exempt it from robots.txt is genuinely contested and has been litigated in public more than once.

The crawling controversy, stated plainly

In 2024 a sequence of investigations, including work published by a major news outlet and later by Cloudflare, alleged that content was being retrieved from addresses and agents not declared in the published crawler list, in cases where the declared agent had been disallowed. Perplexity disputed the characterisation, arguing that user-initiated retrieval and third-party infrastructure sit outside the scope of a crawler policy.

You do not need to resolve who was right to draw the operational lesson. A robots.txt directive is a request honoured voluntarily. If a piece of content genuinely must not be fetched by a machine, the only reliable controls are authentication or not publishing it. For everyone else, the dispute is mostly noise: contractors and service businesses want to be read.

Why Perplexity is the best measurement surface you have

  • Citations are visible by default. You can see exactly which page was used, not merely that a brand was mentioned, which makes attribution unusually clean.

  • Retrieval is live. A page published on Monday can be cited by Friday, so the feedback loop is short enough to test with.

  • Referrals are identifiable. Traffic arrives under the perplexity.ai hostname, which segments cleanly in analytics without any tagging work.

  • Source sets are small. Answers typically lean on a handful of sources, so being in or out is a binary you can log rather than a position you have to interpret.

Use it as the canary. If a change to your page structure improves citation frequency on Perplexity within a fortnight, the same change is very likely helping in slower-moving systems you cannot observe directly.

What to do about it this week

Confirm PerplexityBot is not disallowed, in robots.txt and at the CDN, then check your logs for verified hits rather than assuming the absence of a block means the presence of a crawl. Sites that were never linked from anywhere are frequently uncrawled for reasons that have nothing to do with permissions.

Then run five buyer questions through the product and read the sources it chose. The interesting information is not whether you appeared; it is which competitor page won the passage and why. Nine times out of ten the winning page has a specific number in a short self-contained section, and yours has three paragraphs of context before the answer.

A one-hour Perplexity audit

  1. Fetch your robots.txt and confirm PerplexityBot has no inherited disallow.

  2. Search ninety days of server logs for the agent and record the date of the last verified fetch.

  3. Ask the product your five most commercial buyer questions from a signed-out session.

  4. Open every source it cited instead of you and note the structure of the winning passage.

  5. Rewrite one of your own pages to match that structure, then re-ask the same question a fortnight later.

Want leads like this in your pipeline?

Flock runs the campaigns, screens the enquiries and hands you only the ones that match your service area, job size and capacity. You pay per lead, not per month.

Book a 15-minute fit check  |  See lead package pricing

Related answers

Frequently asked questions

If I block PerplexityBot, am I fully removed from Perplexity?

From the indexed source pool, largely yes. You can still be described using what other sites say about you, and a user pasting your URL directly is a separate path. Blocking removes your ability to be the cited source while leaving third-party descriptions intact.

Does allowing Perplexity mean my content trains a model?

Indexing for retrieval and corpus building for training are different processes with different agents, and vendors document them separately. If training specifically is your concern, handle the training token as its own decision rather than blocking the retrieval crawler that produces your citations.

How quickly does a new page become citable?

Live-retrieval systems move in days rather than months when the page is discoverable. Link it from an indexed page, include it in the sitemap, and expect the first citations within one to three weeks if the content genuinely answers the question better than the incumbent.

Is Perplexity traffic worth anything at this volume?

Judge it on conversion rather than sessions. A visitor arriving after reading a sourced answer that named your pricing and service area has already done the qualification step, which is why the enquiry rate from this channel typically beats generic organic by a wide margin.

NEED A CLEARER PLAN?

Let’s turn your next move into momentum.

Talk to us →

contact

hello@flockleads.com

Reply within 24 hours

REMOTE

Remote-first

Serving clients worldwide

All meetings via Teams or Google Meet

contact

hello@flockleads.com

Reply within 24 hours

REMOTE

Remote-first

Serving clients worldwide

All meetings via Teams or Google Meet

contact

hello@flockleads.com

Reply within 24 hours

REMOTE

Remote-first

Serving clients worldwide

All meetings via Teams or Google Meet