ToolsWebPro logo
Tutorials

Block GPTBot, Keep ChatGPT Search (robots.txt 2026)

Learn how to block OpenAI GPTBot and foundation model training scrapers in robots.txt while allowing OAI-SearchBot and Perplexity for search citations in 2026.

Reviewed by Muhammad Saqlain
·Published 2026-09-28·Updated 2026-09-28·8 min read

Use this guide with

3 ToolsWebPro tools

Open the free tool(s) below and follow the steps in this guide.

Quick answer: The crucial difference between GPTBot and OAI-SearchBot

In 2026, webmasters face a major dilemma when configuring robots.txt: you want to prevent AI companies from scraping your copyrighted articles, source code, and intellectual property to train proprietary foundation models (like GPT-5), but you do not want to sacrifice referral traffic and citations from search engines like ChatGPT Search.

The solution lies in OpenAI's token differentiation: OpenAI operates separate user-agents for different tasks. GPTBot is used exclusively to crawl web content for AI model training. In contrast, OAI-SearchBot and ChatGPT-User are used exclusively by ChatGPT Search to crawl, index, and attribute source links when users search the web in chat.

If you block all OpenAI bots indiscriminately or block User-agent: *, your website becomes invisible in ChatGPT Search citations. By explicitly blocking GPTBot while leaving OAI-SearchBot allowed, you achieve the ideal balance: zero unauthorized training data scraping, with full search visibility and click-through referral traffic.

Direct answer: To block AI model training while preserving ChatGPT Search citations, add 'User-agent: GPTBot' followed by 'Disallow: /' to your robots.txt, then keep 'User-agent: OAI-SearchBot' and 'User-agent: *' allowed. Generate and test your complete file in seconds with our free AI Crawler robots.txt Generator & Validator.

Complete 2026 AI crawler matrix: search vs training bots

Major technology companies (OpenAI, Google, Anthropic, Apple, Meta, ByteDance) now respect distinct User-Agent tokens that differentiate between search citation crawling, model training harvesting, and high-frequency scraping. Understanding this taxonomy is vital for webmaster SEO:

↔ Scroll table horizontally to view full data
Crawler TokenOperatorPrimary FunctionRecommended ActionSEO / Traffic Impact
OAI-SearchBotOpenAIChatGPT Search indexer for web citationsAllowCritical. Allows ChatGPT Search to index your pages and provide clickable citations to users.
ChatGPT-UserOpenAIOn-demand browser fetch when a user pastes your URLAllowEssential. Enables ChatGPT users to analyze, summarize, and reference your web links in real time.
PerplexityBotPerplexity AIPerplexity search engine & conversational citation crawlerAllowCritical for Generative Engine Optimization (GEO). Drives high-intent referral traffic.
ApplebotAppleSiri, Spotlight suggestions, Safari lookups & Apple IntelligenceAllowCritical for iOS/macOS search ecosystem and device-level Siri suggestions.
GPTBotOpenAIHarvests web text to train foundation models (GPT-4/5)DisallowSafe to block. Stops model training. Does NOT hurt ChatGPT Search rankings or citations.
Google-ExtendedGoogleHarvests data to train Gemini & Vertex AI generative modelsDisallowSafe to block. Opts out of Gemini training. Does NOT affect Google Search rankings or Discover.
ClaudeBotAnthropicScrapes web data to train the Claude foundation modelsDisallowSafe to block. Prevents model training while Claude-Web handles on-demand browsing.
Applebot-ExtendedAppleHarvests training datasets for Apple Intelligence generative modelsDisallowSafe to block. Does not block Applebot regular search index or Siri answer links.
Meta-ExternalAgentMetaScrapes web data for training open-source Llama modelsDisallowSafe to block. Stops Meta from training LLMs on your proprietary content.
BytespiderByteDanceAggressive scraper for TikTok and Doubao AI modelsDisallowHighly recommended to block. High-frequency crawls cause severe server CPU load with zero traffic benefit.
CCBotCommon CrawlOpen-source data scraping distributed to commercial AI startupsDisallowSafe to block. Prevents automated extraction and resale of your content in open datasets.
GooglebotGooglePrimary web crawler for organic Google Search indexingAlways AllowNEVER block. Disallowing removes your site completely from Google Search results.

By targeting training crawlers directly, you protect your content library from being ingested into LLM weights without hurting your organic search traffic or AI discovery.

Why blocking Google-Extended does NOT hurt your Google Search rankings

One of the most persistent misconceptions among SEO professionals is that disallowing 'Google-Extended' will damage Google organic rankings or remove the site from Google Search. This is completely false.

Google Search Central has explicitly clarified that Google-Extended is a separate product-control token. Googlebot crawls and indexes pages for Google Search, Google News, and Google Discover. Google-Extended is only used to manage whether a site's content can be utilized to train Gemini, Vertex AI generative APIs, and future model iterations.

Disallowing Google-Extended has zero impact on your core search rankings, crawling frequency, or rich snippets. If you want your site indexed by Google Search but do not want your prose used to teach Gemini how to write, blocking Google-Extended is the official, supported method.

Managing aggressive crawlers: Bytespider and Common Crawl server loads

Beyond copyright and model training concerns, rogue crawlers present a major operational challenge for DevOps and server performance. ByteDance's 'Bytespider' and Common Crawl's 'CCBot' are notorious for executing tens of thousands of requests per hour on un-cached endpoints, dynamic search filters, and product paginations.

Unlike Googlebot or Bingbot, which throttle request rates when server response times rise, Bytespider has frequently been observed ignoring standard Crawl-delay directives on Apache and NGINX servers. Adding an explicit 'Disallow: /' block in robots.txt instructs legitimate spiders to cease crawling. For persistent non-compliant scrapers, you can combine this robots.txt block with Cloudflare WAF or reverse-proxy firewall rules.

Blocking CCBot additionally stops your website from being included in massive public web dumps that third-party AI companies download indiscriminately to train smaller open-source models.

RFC 9309 longest-match evaluation: avoiding path conflicts

In 2022, the Internet Engineering Task Force (IETF) codified the Robots Exclusion Protocol into official internet standard RFC 9309. Prior to RFC 9309, different search engines interpreted conflicting Allow and Disallow directives unpredictably. Understanding the three core rules of RFC 9309 ensures your directives work flawlessly across all modern crawlers:

  • Longest Prefix Wins: Crawlers evaluate all matching path rules and select the one with the greatest number of characters. If you specify 'Disallow: /blog/' (length 7) and 'Allow: /blog/ai-tools' (length 15), a request to /blog/ai-tools is ALLOWED because the Allow rule is longer.
  • Allow Beats Disallow on Equal Length: If both an Allow and a Disallow rule match a path with identical character length (e.g. 'Disallow: /p' and 'Allow: /p'), RFC 9309 dictates that the Allow rule takes precedence.
  • Wildcard (*) and End-of-Line ($) Matching: You can use asterisk (*) to match zero or more characters (e.g. 'Disallow: /*.pdf$') and dollar sign ($) to signify the end of the URL string. This allows surgical blocking of heavy file types without blocking HTML pages.

You can test these complex path matching rules instantly with the live simulator built into our AI Crawler robots.txt Generator. For optimizing agentic LLM discovery alongside robots.txt, ensure you also deploy an audited llms.txt file following our Lighthouse llms.txt guide.

Free AI robots.txt Generator

Generate your custom AI crawler robots.txt in seconds

Pick your crawler strategy, toggle bots, test paths with the RFC 9309 validator, and download your ready-to-deploy robots.txt file.

M

Muhammad Saqlain

Cybersecurity Practitioner & Lead Engineer

Security researcher, web developer, and founder of ToolsWebPro. Tests password entropy, GPU cracking speeds, and client-side encryption systems.

Read full author bio & credentials →