AI Search · Analysis

What blocking Google-Extended, GPTBot or PerplexityBot actually does

Blocking an AI crawler is a one-line change in robots.txt, but each line controls something different. Here is what Google, OpenAI and Perplexity document for each token, and where the documentation says nothing.

By Dean Cruddace · Published · Updated · Read 6 min

Key findings

  • Google says Google-Extended does not affect a site's inclusion in Google Search and is not a ranking signal. It covers use of crawled content for Gemini training and for grounding in Gemini Apps and Vertex AI.
  • OpenAI documents GPTBot (training) and OAI-SearchBot (search) as independent settings, and says sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers.
  • Perplexity says PerplexityBot surfaces and links sites in its search results and is not used to crawl for AI foundation models, while Perplexity-User generally ignores robots.txt.

Many site owners are asking whether to block AI crawlers, and the advice they find tends to treat the bots as one group. The vendors’ own documentation does not. Each crawler or token is described as doing a specific job, and the effect of blocking it differs. This article sets out what Google, OpenAI and Perplexity say in their documentation, as fetched in October 2026, and is explicit about what those pages do not say. Whether to block any of them is a decision for the site owner; our aim is to make sure it is made on what is documented.

Robots.txt asks; it does not enforce

All three vendors use robots.txt as the mechanism, so its limits apply. Google’s introduction to robots.txt says the file tells crawlers which URLs they can access, and that it is mainly for managing crawler traffic, not a mechanism for keeping a page out of Google. It also says the instructions cannot enforce crawler behaviour: it is up to the crawler to obey them. Google says respectable crawlers do, while others might not, and that different crawlers may interpret rules differently.

That framing matters for the rest of this article. A robots.txt rule is a request that a named crawler can honour. It is not access control, and for anything confidential the same page recommends other methods such as password protection.

Google-Extended: a token, not a crawler

Google’s list of common crawlers describes Google-Extended as a standalone product token. It has no separate HTTP user agent string: crawling is done with existing Google user agent strings, and the robots.txt token is used in a control capacity. Publishers can use it to manage whether content Google crawls may be used for training future generations of Gemini models, and for grounding (providing content from the Google Search index to the model at prompt time) in Gemini Apps and in Grounding with Google Search on Vertex AI.

The page states what it does not do: Google-Extended does not impact a site’s inclusion in Google Search, nor is it used as a ranking signal. Google’s AI features guide is consistent with that. It says robots.txt directives for Googlebot are the control for how a site is crawled for Search, and points to Google-Extended only for limiting AI training and grounding in some of Google’s other systems. In our reading, blocking Google-Extended is therefore not a way to opt out of AI Overviews or AI Mode, and the documentation does not say it is.

GPTBot and OAI-SearchBot: two independent settings

OpenAI’s crawler overview says it uses OAI-SearchBot and GPTBot robots.txt tags so webmasters can manage how their content works with AI, and that each setting is independent of the others. Its own example is a site that allows OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training its generative AI foundation models.

GPTBot is described as crawling content that may be used in training, and OpenAI says disallowing it indicates a site’s content should not be used for that. OAI-SearchBot is described as the crawler used to surface websites in ChatGPT’s search features. OpenAI says sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links, and it recommends allowing it. A hypothetical consequence: a site that disallows both would be opting out of training and of ChatGPT search answers, and a site that disallows only GPTBot would not be opting out of search. That follows from OpenAI’s wording rather than from anything it states about a particular site.

A third agent matters for logs. OpenAI describes ChatGPT-User as used for certain user actions in ChatGPT and Custom GPTs, says it is not used for automatic crawling, and says that because these actions are user-initiated, robots.txt rules may not apply. It says OAI-SearchBot, not ChatGPT-User, is the one to use in robots.txt for search opt-outs.

PerplexityBot and Perplexity-User

Perplexity’s crawler documentation describes PerplexityBot as designed to surface and link websites in search results on Perplexity, and says it is not used to crawl content for AI foundation models. It recommends allowing PerplexityBot if you want your site to appear in those results. On the page we read, Perplexity does not list a separate training crawler token.

Perplexity-User is different. Perplexity says it supports user actions within Perplexity: when a user asks a question it might visit a page and include a link in its response. It is not used for web crawling or to collect training content, and because a user requested the fetch, Perplexity says this fetcher generally ignores robots.txt rules. Perplexity also says that if you use a web application firewall you may need to explicitly allow its bots, and publishes IP ranges for that purpose. In other words, a block applied at the firewall and a block applied in robots.txt are different mechanisms with different effects.

Where the documentation is silent

Several questions a site owner might reasonably ask are not answered on these pages. We found nothing that says whether blocking a token removes content already collected, or whether it changes what a model already trained on that content says. None of the pages says that allowing a crawler leads to a mention or a citation, and none says that blocking one affects traffic. Timing is documented only loosely: OpenAI says it can take around 24 hours for its search systems to adjust after a robots.txt update, and Perplexity says changes may take up to 24 hours.

The pages also cover only these vendors. Other AI companies run their own crawlers with their own documentation, and we have not covered them here. Anything not stated above should be treated as unknown, not as a hidden effect.

Where this fits at Cultured Digital: access as a decision

We treat crawler access as a technical-access question inside our AI & Search Visibility work. Technical access asks whether crawlers can reach and read the content; the machine understanding guide covers sites optimised for machine reading, and robots.txt is part of our crawling and indexing work in technical SEO. How AI tools describe a brand is looked at separately, in the visibility review.

Blocking or allowing is the site owner’s decision. Our part is to set out what each choice is documented to do, and to say plainly when the documentation stops. We cannot promise that any setting produces citations.

Further reading on AI crawler controls

Written by Dean Cruddace

Founder of Cultured Digital. Working in SEO since 2001, across independent consultancy, in-house and agency roles, with a focus on technical SEO, strategy and development.

About Dean →