Free to explore: filter winning sites by DR, traffic and niche  ·  Try the live explorer →

How to block AI crawlers with robots.txt

A robot hand near a laptop keyboard

Most guides to block ai crawlers robots txt rules hand you a long list of user agents and a blanket Disallow: / for each. That is the wrong starting point. The useful question is not “how do I stop AI bots?” but “which uses of my content am I refusing?” Training a model and citing your page in a search answer are different transactions, and the main operators now let you accept one and refuse the other.

This post works through what OpenAI and Google actually document, what each rule does, and the trade-off most publishers miss. If you need a refresher on the file itself, our guide to robots.txt for SEO covers syntax and precedence.

Which AI crawlers you can block, and what each one does

OpenAI’s overview of its crawlers lists three user agents with three distinct jobs. GPTBot crawls content that may be used to train generative AI foundation models; disallowing it signals that your content should not be used for training. OAI-SearchBot is the bot that surfaces sites in ChatGPT search. ChatGPT-User covers actions a user starts, such as asking ChatGPT to visit a page.

A robots file with user-agent lines

Google takes a different route. Google-Extended, described on Google Search Central’s common crawlers page, controls whether content “may be used for training future generations of Gemini models” and for grounding. It is a robots.txt token, not a separate bot: it “doesn’t have a separate HTTP request user agent string”. Crawling still happens through Googlebot; the token only governs the AI use.

TokenOperatorWhat it controlsEffect of blocking
GPTBotOpenAICrawling for training foundation modelsSignals the content shouldn’t be used for training
OAI-SearchBotOpenAISurfacing sites in ChatGPT search“will not be shown in ChatGPT search answers”
ChatGPT-UserOpenAIUser-initiated actions“robots.txt rules may not apply”
Google-ExtendedGoogleGemini training and grounding“does not impact a site’s inclusion in Google Search”

Robots.txt rules for GPTBot and Google-Extended

The rule itself is ordinary robots.txt. To refuse OpenAI’s training crawler across the whole site, add a group for its user agent:

The settings are independent, so you can allow search while refusing training. A file that keeps you eligible for ChatGPT search, opts out of OpenAI training and opts out of Gemini training looks like this:

Bot visits in a server log

You can also scope a block to a directory instead of the whole site, for example Disallow: /members/ under GPTBot if only gated material worries you. Remember that a crawler follows the most specific group that names it, so a User-agent: * block does not override a named group further down.

Two practical details trip people up. The token has to be exactly the one the operator documents, so copy it from the source page rather than from a forum list that may be out of date. And a disallow for a bot you have never seen in your logs costs nothing, while a disallow for one that sends you visitors costs a great deal. Write the file with that asymmetry in mind.

Changes are not instant. OpenAI says “it can take ~24 hours from a site’s robots.txt update for our systems to adjust.” Edit the file, check it serves at https://example.com/robots.txt, then give it a day before checking your logs for the result.

Do AI crawlers respect robots.txt?

The documented ones say they do, with one stated exception. For ChatGPT-User, OpenAI writes that “robots.txt rules may not apply”, because the fetch is an action a person asked for rather than an automated crawl. Blocking ChatGPT-User in robots.txt is therefore not a reliable way to stop those requests.

A chatbot answer citing sources

More generally, robots.txt is a voluntary convention. It tells well-behaved crawlers what you want; it does not technically stop anything. A scraper that ignores it, or that does not identify itself honestly, will not be deterred by a line in a text file. If you need enforcement rather than a request, that happens at the server or CDN level, through bot management or IP rules, and it sits outside what robots.txt can do.

It is also worth separating robots.txt from llms.txt, which often gets mentioned in the same breath. Our explainer on what llms.txt is covers why it is a suggestion file for models, not an access control.

The trade-off: blocking AI crawlers can cost you AI visibility

This is where blanket blocks go wrong. OpenAI is explicit: “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers”, though they may still appear as navigational links. Copy a list of “AI bots” into your file and you may be refusing citations, not just training.

Bot management settings in a CDN

Received wisdom says blocking AI is the protective choice. For many sites it is the opposite. If AI answers are becoming a discovery channel for your topic, the bot that puts you in those answers is the one you want crawling. Our guides on how to rank in ChatGPT and GEO versus SEO cover what earning those citations involves.

Google has removed the dilemma on its side. Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” Opting out of Gemini training costs you nothing in organic rankings, which makes it the lowest-risk block on this list.

A publisher weighing traffic sources

A sensible default for most sites

Decide per use, not per company. Refusing training is a reasonable position if you sell content, data or expertise you do not want absorbed into a model. Refusing search visibility is a different decision and should be a deliberate one. For most publishers that leaves a clear default: block GPTBot and Google-Extended if training bothers you, allow OAI-SearchBot, and leave Googlebot alone.

A developer editing a configuration file

Agencies and multi-site owners should make the choice once per site type rather than once per site. A reference publisher that lives on citations has more to lose from a search-bot block than a membership site whose value sits behind a login. Write the policy down, apply it consistently, and revisit it when an operator documents a new user agent.

Then verify. Fetch your live robots.txt, confirm each group names the right token, and watch your server logs over the following days for the user agents you blocked. If a bot keeps arriving after the adjustment window, robots.txt has done all it can, and the next step is a server-side rule.

One limitation to state plainly. Robots.txt only governs future crawling by bots that choose to honour it. Neither OpenAI’s nor Google’s documentation, as summarised here, says that blocking removes content already collected, and nothing in robots.txt lets you confirm what any operator has stored. Treat a block as a forward-looking instruction, not a deletion request.

Frequently asked questions

Does blocking GPTBot remove me from ChatGPT search?

No. ChatGPT search uses OAI-SearchBot, and OpenAI says the settings are independent. You can block GPTBot and still allow OAI-SearchBot.

Does blocking Google-Extended hurt my Google rankings?

Google says it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.”

Do AI crawlers respect robots.txt?

The documented crawlers say they do, but OpenAI notes that for user-initiated ChatGPT-User requests “robots.txt rules may not apply”. Robots.txt is voluntary, so enforcement needs server or CDN rules.

How long does a robots.txt change take to apply?

For OpenAI’s crawlers, “it can take ~24 hours from a site’s robots.txt update for our systems to adjust.”

The takeaway Block by purpose, not by brand. GPTBot and Google-Extended govern training; OAI-SearchBot governs whether ChatGPT search shows you. Refuse training if you want to, but block the search bot only if you are willing to disappear from those answers.