THE SHORT ANSWER
Google-Extended is a control token you place in robots.txt, not a crawler that visits your site. It governs whether content Googlebot already fetched may be used for Gemini model training and certain grounding uses. Crucially it does not remove you from AI Overviews or AI Mode, which run on the ordinary Search index. The only levers there are snippet directives, and those cost you your normal search snippets too.
This is the single most misunderstood setting in the entire AI-crawler discussion, and the misunderstanding is expensive in both directions. Publishers add the token expecting to disappear from generative results in Search and are baffled when nothing changes. Marketers avoid it expecting a ranking penalty that does not exist. Both then plan around the wrong model for years.
The confusion is understandable, because Google made an unusual design choice. It separated the training permission from the search permission while leaving both attached to one crawl by one agent, so the file you edit and the behaviour you observe belong to different systems. Once that split is clear in your head, every practical consequence on this page follows from it.
The numbers, at a glance
What it is: a robots.txt user-agent token with no crawler behind it; Googlebot does the fetching either way
What it controls: use of your content for Gemini model training and for grounding in Gemini surfaces
What it does not control: AI Overviews and AI Mode in Search, which are governed by ordinary Search indexing
The only Search-side lever: nosnippet, max-snippet and data-nosnippet, all of which also suppress your normal result snippet
The design decision that causes all the confusion
Googlebot crawls your site once. The resulting content then flows to several consumers: the Search index, the generative features built on top of that index, and the model-training pipeline. Google-Extended is a valve on the last of those, expressed through the robots.txt vocabulary purely because that file is where site owners already look. No request is ever made by an agent calling itself Google-Extended, which is why it never appears in your access logs.
Because the generative features in Search read from the Search index, opting out of training leaves them entirely unaffected. Google has been consistent that participation in AI Overviews follows from being indexed and snippet-eligible, not from any separate opt-in. There is no button, and any tool claiming to provide one is selling something else.
So how do you actually stay out of AI Overviews?
nosnippet. Prevents any text snippet being shown, which also prevents the passage being used generatively. It removes your descriptive snippet from ordinary results at the same time.
max-snippet with a low value. Caps the characters available. Setting it to zero behaves like nosnippet; small non-zero values limit what can be lifted while keeping some text.
data-nosnippet. An HTML attribute that excludes a specific element rather than the whole page, which is the only surgical option in the set.
noindex. Removes the page from Search entirely, generative features included, at the obvious cost.
Every option on that list is a trade, and for a service business every one of them is a bad trade. You would be suppressing the description that persuades a buyer in order to avoid being quoted to that same buyer. The publisher calculus is different because their snippet substitutes for a visit they were monetising.
The related tokens people confuse with it
Google-CloudVertexBot fetches sites at the request of a Vertex AI customer building a grounded application, which is an enterprise product path rather than consumer Search. Googlebot-News, Google-InspectionTool and the various media crawlers are separate again. Treating the whole family as one policy question is how sites end up blocking something they meant to keep.
There is also a naming trap: because Google-Extended is a token and not an agent, a wildcard disallow written against all user-agents does apply to it. Plenty of sites have opted out of Gemini training without ever intending to, through a staging rule that survived launch.
What most service businesses should actually set
Leave Google-Extended permissive unless you have a licensing strategy that depends on withholding. The upside of the block is theoretical and future-dated; the downside is that you have voluntarily removed yourself from a system that increasingly answers questions about which contractor to call, in exchange for nothing you can bank.
Spend the attention you were about to spend on this token on being snippet-eligible and extractable instead. A page with a clear forty-word answer near the top, a specific figure and a heading that matches the question is doing more for your generative visibility than any robots directive can.
Getting the Google-side controls right
Search robots.txt for a wildcard disallow that catches Google-Extended by accident.
Decide the training question explicitly and write the token in either way, so the intent is documented.
Audit page templates for a stray nosnippet or low max-snippet left over from a previous project.
Reserve data-nosnippet for the handful of elements that genuinely should not be quoted out of context.
Confirm in Search Console that your key pages are indexed, since indexing is the real precondition.
Want leads like this in your pipeline?
Flock runs the campaigns, screens the enquiries and hands you only the ones that match your service area, job size and capacity. You pay per lead, not per month.
Book a 15-minute fit check | See lead package pricing
Related answers
Frequently asked questions
Does adding Google-Extended hurt my rankings?
No. It is a permission token consumed downstream of crawling and it carries no ranking signal. Classic Search behaviour, indexing and result placement are unaffected. The only thing it changes is whether the content may be used for model training and certain Gemini grounding uses.
Can I appear in AI Overviews but not in Gemini's training data?
Yes, and that is the configuration the split was designed to allow. Disallow the training token while remaining indexed and snippet-eligible. You stay available to the generative features built on Search and stay out of the training corpus, which is precisely the trade most publishers were asking for when the token first appeared.
Why does Google-Extended never appear in my logs?
Because nothing fetches under that name. It is a directive interpreted after Googlebot has already retrieved the page. Expecting it in access logs is the clearest sign someone has modelled it as a crawler rather than a permission, which is the misconception behind most of the bad advice on this subject.
Is there any way to opt out of AI Overviews specifically?
Not as a discrete setting. The available controls are snippet-level and blunt, and they degrade your ordinary search presence at the same time. Google has consistently declined to offer a generative-only opt-out, and there is no sign of that changing.
NEED A CLEARER PLAN?
Let’s turn your next move into momentum.
Talk to us →