How should robots.txt be configured for AI crawlers in 2026?

How should robots.txt be configured for AI crawlers in 2026?

How should robots.txt be configured for AI crawlers in 2026?

THE SHORT ANSWER

Handle each AI crawler explicitly rather than with a blanket rule, because the agents serve different purposes and deserve different answers. For most service businesses the right configuration allows the retrieval and answer crawlers, since those produce citations, and treats training-only crawlers as a separate decision. Then check your CDN bot protection, which blocks AI agents on several platforms regardless of what robots.txt says.

Robots.txt was designed in 1994 for a web with one kind of crawler. It is now carrying a policy decision it was never built for, which is why most sites have an inherited configuration nobody has read in three years.

The fix is not complicated, but it does require deciding rather than defaulting.

The numbers, at a glance

  • Agents worth naming individually: GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, Bingbot, CCBot

  • Most common mistake: a blanket disallow inherited from a staging configuration or a previous agency

  • Second most common: CDN bot protection blocking agents that robots.txt explicitly allows

  • Directive that does not exist: there is no standard way to allow crawling while refusing training; it varies per vendor

The agents worth handling individually

  • Training crawlers fetch pages to build the corpus a model learns from. Blocking them affects what the model knows in future versions, months out.

  • Retrieval and search crawlers build an index the assistant queries when it needs a current answer. Blocking them affects citations now.

  • Live user-triggered fetchers retrieve a specific page because a user asked. Many operators treat these as user agents rather than crawlers, and they may not follow robots.txt at all.

  • Common Crawl is an open dataset many models train on indirectly. Blocking it affects a wide set of downstream systems at once.

The practical implication is that a single allow-or-block decision is too coarse. Most service businesses want the retrieval crawlers unambiguously, are indifferent to the training crawlers, and should decide about Common Crawl on the same basis as the training ones.

The configuration most service businesses should run

Allow the retrieval and search agents without qualification, since they are the mechanism by which a current, accurate description of your business reaches a buyer mid-decision. Allow the training crawlers unless your content is your product. Keep any disallow rules narrow and path-specific rather than domain-wide.

Then add the boring but decisive step: verify. Fetch your own robots.txt from a clean browser session and read every line. Sites that migrated platforms, changed agencies or came out of a staging environment routinely carry a blanket disallow that nobody notices for months, because nothing visibly breaks.

The CDN problem that overrides everything above

Bot protection on Cloudflare, Akamai, AWS WAF and several hosting platforms includes AI-crawler categories that can be switched on by default or by a well-meaning administrator. Those rules operate at the network layer and reject the request before robots.txt is ever read.

The symptom is specific and easy to recognise: your robots.txt is permissive, your content is good, and yet no assistant can quote a page it has clearly never fetched. Check the bot-management settings directly rather than inferring from robots.txt, and if your platform offers a verified-bot allowlist, put the AI agents on it explicitly.

What robots.txt cannot do for you

It cannot remove knowledge a model already has. It cannot stop a user pasting your page into a chat window. It cannot express "you may read this but not train on it" in any standardised way. And it cannot be relied on by agents that classify themselves as user-initiated rather than automated.

If a piece of content genuinely must not be repeated by a machine, the only reliable control is not publishing it openly. Authentication works; a robots directive is a request, honoured by reputable operators and ignored by the rest.

Getting this right in one sitting

  1. Fetch and read your live robots.txt end to end.

  2. Name each AI agent explicitly instead of relying on a wildcard.

  3. Remove any inherited blanket disallow left over from staging or a migration.

  4. Open the CDN bot-management panel and check the AI-crawler category.

  5. Re-verify after every platform migration, without exception.

Want leads like this in your pipeline?

Flock runs the campaigns, screens the enquiries and hands you only the ones that match your service area, job size and capacity. You pay per lead, not per month.

Book a 15-minute fit check  |  See lead package pricing

Related answers

Frequently asked questions

Does a wildcard user-agent rule cover AI crawlers?

Usually yes, which is precisely the problem: a wildcard disallow written for a staging site blocks every AI agent silently. Name the agents you care about explicitly so the decision is visible in the file.

Will allowing AI crawlers slow my site down?

Their crawl volume is modest against ordinary search crawling for a site of typical size. If crawl load is genuinely a problem, rate-limit rather than block.

Should I add an llms.txt file as well?

It costs almost nothing and a growing number of tools read it. Treat it as a supplement to a correct robots.txt and a real sitemap, not a replacement for either.

How often should this be reviewed?

Twice a year, and immediately after any platform migration, CDN change or agency handover. Those three events cause nearly every robots.txt regression.

NEED A CLEARER PLAN?

Let’s turn your next move into momentum.

Talk to us →

contact

hello@flockleads.com

Reply within 24 hours

REMOTE

Remote-first

Serving clients worldwide

All meetings via Teams or Google Meet

contact

hello@flockleads.com

Reply within 24 hours

REMOTE

Remote-first

Serving clients worldwide

All meetings via Teams or Google Meet

contact

hello@flockleads.com

Reply within 24 hours

REMOTE

Remote-first

Serving clients worldwide

All meetings via Teams or Google Meet