Should I allow GPTBot to crawl my site?

Should I allow GPTBot to crawl my site?

Should I allow GPTBot to crawl my site?

THE SHORT ANSWER

For a service business that wants to be found and recommended, allow it. Blocking GPTBot removes your pages from the pool a model can read and cite, and protects nothing you were monetising. The block only makes sense if your content is the product you sell, which is the publisher case, not the contractor case. Note that GPTBot and the live browsing fetcher are separate agents, so blocking one does not block the other.

The default advice circulating in 2024 was to block everything. It came from publishers with a genuine grievance about training data and it was copied wholesale by businesses with the opposite interest.

The question worth asking is not whether the model should be allowed to learn from you. It is whether you would rather be described accurately or not described at all.

The numbers, at a glance

  • What GPTBot does: crawls pages for model training and, in some configurations, for retrieval

  • What blocking achieves: removes you from the source pool; does not remove existing knowledge already trained in

  • Who should block: publishers and paywalled archives whose content is the product

  • Who should allow: service businesses, agencies, contractors and anyone whose content is marketing

What GPTBot is, precisely

GPTBot is the crawler that fetches public web pages on behalf of OpenAI. It identifies itself in the user-agent string and respects robots.txt directives in the ordinary way. It is not the same agent as the live browsing fetch that runs when a user asks a question and the assistant goes to look something up in real time.

This distinction causes most of the confusion. A site can block GPTBot and still be fetched during a live browsing session, and a site can allow GPTBot while a separate agent is blocked by a firewall rule nobody remembers writing. Check what you are actually blocking before concluding anything about why you are not being cited.

The case for allowing it, for a service business

Your website exists to be read. Every euro spent on it was spent to make information about your service available to people evaluating you. An assistant reading that information and repeating it to a buyer is the same transaction as a human reading it, at lower cost to you.

The specific risk of blocking is not zero visibility, it is inaccurate visibility. Models still know your company exists from directories, review sites, press mentions and competitors' comparison pages. Blocking your own site removes the one source that is under your control, leaving the assistant to describe you using other people's descriptions, which are older, thinner and occasionally wrong.

The case for blocking, and who it genuinely applies to

If your content is the thing customers pay for, allowing free extraction into a system that answers questions without sending traffic is a real transfer of value. News organisations, research firms, course providers and subscription publications have a coherent reason to block or to license instead.

There is also a narrower legitimate case: content you do not want repeated out of context, such as detailed pricing that only makes sense with conditions attached, or client work published under an agreement that limits reuse. Handle that at page level with a targeted disallow rather than blocking the whole domain, which is a much larger decision than the problem requires.

How to implement whichever choice you make

  • To allow: do nothing, as long as robots.txt does not contain an inherited blanket disallow. Confirm rather than assume; many sites carry blocks added by a previous agency.

  • To block entirely: add a user-agent block for GPTBot with a disallow of the root path.

  • To block selectively: allow the root and disallow only the specific paths you want withheld.

  • Either way: check your CDN and firewall as well. Bot-protection rules block AI crawlers by default on several platforms, silently, regardless of what robots.txt says.

A ten-minute audit

  1. Fetch your own robots.txt and read it, rather than trusting what you think it says.

  2. List every AI user-agent it mentions and decide, deliberately, on each one.

  3. Check the CDN or firewall bot rules for a blanket AI-crawler block.

  4. Confirm the pages you most want cited are not caught by an unrelated disallow.

  5. Re-check after any site migration, since robots.txt is routinely lost in one.

Want leads like this in your pipeline?

Flock runs the campaigns, screens the enquiries and hands you only the ones that match your service area, job size and capacity. You pay per lead, not per month.

Book a 15-minute fit check  |  See lead package pricing

Related answers

Frequently asked questions

If I block GPTBot, will ChatGPT stop mentioning my company?

No. It will keep whatever it learned previously and whatever it can find on other sites. You will have removed your own site as a correction mechanism, which usually makes the description worse rather than absent.

Does blocking GPTBot affect Google rankings?

No. It is a separate crawler from Googlebot and has no bearing on classic ranking or on Google's AI Overviews, which use Google's own infrastructure.

Can I allow crawling but refuse training?

Some vendors separate those permissions with distinct user-agents, which is why the crawler list is worth handling individually rather than with one blanket rule. The controls differ per vendor and change, so re-read the current documentation before relying on a specific directive.

Is there any benefit to blocking for a contractor?

Almost never. The realistic outcome is that assistants describe your business using directory listings and competitor comparison pages instead of your own site.

NEED A CLEARER PLAN?

Let’s turn your next move into momentum.

Talk to us →

contact

hello@flockleads.com

Reply within 24 hours

REMOTE

Remote-first

Serving clients worldwide

All meetings via Teams or Google Meet

contact

hello@flockleads.com

Reply within 24 hours

REMOTE

Remote-first

Serving clients worldwide

All meetings via Teams or Google Meet

contact

hello@flockleads.com

Reply within 24 hours

REMOTE

Remote-first

Serving clients worldwide

All meetings via Teams or Google Meet