Learn · AI Search · Beginner

How to decide which AI crawlers to allow or block

Learn what Google-Extended, GPTBot, OAI-SearchBot and PerplexityBot are each documented to do, then write, upload and check a robots.txt file that reflects your choice.

By Dean Cruddace · 45 minutes to do · Updated · Last reviewed

What AI crawlers are

An AI crawler is an automated program that fetches web pages for an AI company's products. Each company publishes a short name for its crawlers, and you can use that name in a plain text file called robots.txt to ask it to stay away from some or all of your site.

The names are not one group. Each company documents separate jobs for separate names, such as training a model or finding pages for a search answer. Blocking one does not block the others.

Why the allow-or-block decision matters

Google says a robots.txt file tells crawlers which URLs they can access and that it cannot enforce behaviour: it is up to the crawler to obey (Google's robots.txt introduction). So a rule is a request, and you need to know what you are requesting. OpenAI, for example, says its crawler settings are independent of each other, so a site can allow one for search results while disallowing another to indicate its content should not be used for training (OpenAI's crawler overview). Choosing without reading this can block something you wanted, or leave open something you did not.

What you need to decide which AI crawlers to allow or block

  • The ability to view and edit your site's robots.txt file, which on a site builder or CMS may be a setting or plugin
  • A plain text editor such as Notepad (not a word processor)
  • Access to Google Search Console for the site

Decide which AI crawlers to allow or block, step by step

  1. 1

    Look at your current robots.txt

    Type your site's address followed by /robots.txt into the browser. Google says the file must sit at the root of the site and applies only to the protocol, host and port it is posted on, so a subdomain needs its own. Copy what you see into a note. On a hosted builder you may not be able to edit it directly. See how to create a robots.txt file.

    You will know it worked when You have a saved copy of the current file, or you know there is none (Google says that without rules, everything is implicitly allowed to be crawled).

  2. 2

    List what each name is documented to do

    Write one line per name, from the vendors' own pages. Google-Extended: manages whether crawled content may be used for training future Gemini models and for grounding in Gemini Apps and Vertex AI; Google says it does not affect inclusion in Google Search (Google's common crawlers). GPTBot: crawls content that may be used for training. OAI-SearchBot: surfaces sites in ChatGPT's search features. ChatGPT-User: acts on user requests; OpenAI says robots.txt rules may not apply. PerplexityBot: surfaces and links sites in Perplexity search results and is not used for foundation model training. Perplexity-User: generally ignores robots.txt.

    You will know it worked when You have a list of names, each with one documented purpose and a note of whether robots.txt controls it.

  3. 3

    Decide per name and record why

    For each name robots.txt can control, choose allow or block. OpenAI says sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, and recommends allowing it to appear in search results; Perplexity recommends allowing PerplexityBot for the same reason. Neither page says that allowing or blocking affects traffic. Write down each decision and the reason.

    You will know it worked when Each name has a decision and a one-line reason, agreed by whoever owns the decision.

  4. 4

    Write the rules

    In a plain text editor, add one group per crawler: a User-agent: line naming it, then a Disallow: line. Google gives Disallow: / as the way to block a named agent from the whole site. Hypothetical example, blocking only training use by two vendors:

    User-agent: GPTBot
    Disallow: /

    User-agent: Google-Extended
    Disallow: /

    Google says rules are case-sensitive and the file should be saved as UTF-8.

    You will know it worked when Reading the file back, every group begins with a User-agent line and the names match the vendors' spelling.

  5. 5

    Upload it and open it in a private window

    Put the file at the root of your site, or change the setting in your CMS. Then open a private browsing window and visit /robots.txt, as Google advises. OpenAI says its search systems can take around 24 hours to adjust; Perplexity says up to 24 hours.

    You will know it worked when The private window shows your new rules at the root address.

  6. 6

    Check Google can read the file

    Google says the robots.txt report in Search Console helps fix issues with markup, for files already accessible on your site. Open it and look for problems. This confirms Google can parse the file; the pages we read do not say it tests other companies' crawlers.

    You will know it worked when The report shows no errors for your file.

Common mistakes when you decide which AI crawlers to allow or block

  • Expecting Google-Extended to remove you from AI Overviews or AI Mode. Google says robots.txt directives for Googlebot are the control for crawling for Search.
  • Treating robots.txt as security. Google says it cannot enforce behaviour.
  • Blocking OAI-SearchBot or PerplexityBot and still expecting to appear in ChatGPT or Perplexity search answers.
  • Assuming a block reaches user-triggered fetchers such as ChatGPT-User or Perplexity-User.
  • Assuming a block undoes the past. The pages we read do not say that it removes content already collected.

Terms used when you decide which AI crawlers to allow or block

robots.txt
A plain text file at the root of a site that asks crawlers which URLs they may fetch.
User-agent token
The short name a crawler answers to in robots.txt, such as GPTBot.
Grounding
Google's term for providing content to an AI model at the moment it answers, to improve factual accuracy.
Training
Using collected content to build or improve an AI model, as opposed to showing a page in an answer.
Web application firewall
A security layer in front of a site. Perplexity says you may need to allow its bots there, separately from robots.txt.
Unsure what to allow in robots.txt?

Send Dean your current robots.txt and what you want to protect or be found for. Access decisions sit between technical SEO and AI search visibility.

Talk to an SEO specialistAI search visibility: machine understanding →