Block GPTBot, Keep ChatGPT Search (robots.txt 2026)
Learn how to block OpenAI GPTBot and foundation model training scrapers in robots.txt while allowing OAI-SearchBot and Perplexity for search citations in 2026.
Use this guide with
3 ToolsWebPro tools
Open the free tool(s) below and follow the steps in this guide.
- •Quick answer: The crucial difference between GPTBot and OAI-SearchBot
- •Complete 2026 AI crawler matrix: search vs training bots
- •The recommended 2026 robots.txt snippet (Copy-paste ready)
- •Why blocking Google-Extended does NOT hurt your Google Search rankings
- •Managing aggressive crawlers: Bytespider and Common Crawl server loads
- •RFC 9309 longest-match evaluation: avoiding path conflicts
Quick answer: The crucial difference between GPTBot and OAI-SearchBot
In 2026, webmasters face a major dilemma when configuring robots.txt: you want to prevent AI companies from scraping your copyrighted articles, source code, and intellectual property to train proprietary foundation models (like GPT-5), but you do not want to sacrifice referral traffic and citations from search engines like ChatGPT Search.
The solution lies in OpenAI's token differentiation: OpenAI operates separate user-agents for different tasks. GPTBot is used exclusively to crawl web content for AI model training. In contrast, OAI-SearchBot and ChatGPT-User are used exclusively by ChatGPT Search to crawl, index, and attribute source links when users search the web in chat.
If you block all OpenAI bots indiscriminately or block User-agent: *, your website becomes invisible in ChatGPT Search citations. By explicitly blocking GPTBot while leaving OAI-SearchBot allowed, you achieve the ideal balance: zero unauthorized training data scraping, with full search visibility and click-through referral traffic.
Direct answer: To block AI model training while preserving ChatGPT Search citations, add 'User-agent: GPTBot' followed by 'Disallow: /' to your robots.txt, then keep 'User-agent: OAI-SearchBot' and 'User-agent: *' allowed. Generate and test your complete file in seconds with our free AI Crawler robots.txt Generator & Validator.
Open the tool:
Complete 2026 AI crawler matrix: search vs training bots
Major technology companies (OpenAI, Google, Anthropic, Apple, Meta, ByteDance) now respect distinct User-Agent tokens that differentiate between search citation crawling, model training harvesting, and high-frequency scraping. Understanding this taxonomy is vital for webmaster SEO:
| Crawler Token | Operator | Primary Function | Recommended Action | SEO / Traffic Impact |
|---|---|---|---|---|
| OAI-SearchBot | OpenAI | ChatGPT Search indexer for web citations | Allow | Critical. Allows ChatGPT Search to index your pages and provide clickable citations to users. |
| ChatGPT-User | OpenAI | On-demand browser fetch when a user pastes your URL | Allow | Essential. Enables ChatGPT users to analyze, summarize, and reference your web links in real time. |
| PerplexityBot | Perplexity AI | Perplexity search engine & conversational citation crawler | Allow | Critical for Generative Engine Optimization (GEO). Drives high-intent referral traffic. |
| Applebot | Apple | Siri, Spotlight suggestions, Safari lookups & Apple Intelligence | Allow | Critical for iOS/macOS search ecosystem and device-level Siri suggestions. |
| GPTBot | OpenAI | Harvests web text to train foundation models (GPT-4/5) | Disallow | Safe to block. Stops model training. Does NOT hurt ChatGPT Search rankings or citations. |
| Google-Extended | Harvests data to train Gemini & Vertex AI generative models | Disallow | Safe to block. Opts out of Gemini training. Does NOT affect Google Search rankings or Discover. | |
| ClaudeBot | Anthropic | Scrapes web data to train the Claude foundation models | Disallow | Safe to block. Prevents model training while Claude-Web handles on-demand browsing. |
| Applebot-Extended | Apple | Harvests training datasets for Apple Intelligence generative models | Disallow | Safe to block. Does not block Applebot regular search index or Siri answer links. |
| Meta-ExternalAgent | Meta | Scrapes web data for training open-source Llama models | Disallow | Safe to block. Stops Meta from training LLMs on your proprietary content. |
| Bytespider | ByteDance | Aggressive scraper for TikTok and Doubao AI models | Disallow | Highly recommended to block. High-frequency crawls cause severe server CPU load with zero traffic benefit. |
| CCBot | Common Crawl | Open-source data scraping distributed to commercial AI startups | Disallow | Safe to block. Prevents automated extraction and resale of your content in open datasets. |
| Googlebot | Primary web crawler for organic Google Search indexing | Always Allow | NEVER block. Disallowing removes your site completely from Google Search results. |
By targeting training crawlers directly, you protect your content library from being ingested into LLM weights without hurting your organic search traffic or AI discovery.
Open the tool:
The recommended 2026 robots.txt snippet (Copy-paste ready)
Here is the standard-compliant robots.txt template recommended for modern content creators, SaaS platforms, and digital publishers in 2026. It blocks commercial AI model training scrapers and high-load crawlers while allowing Googlebot, Bingbot, and ChatGPT Search:
# Block AI Model Training Scrapers User-agent: GPTBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Meta-ExternalFetcher Disallow: / User-agent: cohere-ai Disallow: / # Block High-Frequency Commercial Scrapers User-agent: Bytespider Disallow: / User-agent: CCBot Disallow: / User-agent: Amazonbot Disallow: / User-agent: Diffbot Disallow: / # Allow All Other Bots (Search Engines & AI Search Citations) User-agent: * Disallow: /admin/ Disallow: /private/ Disallow: /api/ Sitemap: https://yourdomain.com/sitemap.xml
Notice how OAI-SearchBot, PerplexityBot, Applebot, Googlebot, and Bingbot are NOT listed in the disallowed user-agent blocks. Under standard Robots Exclusion Protocol rules, any crawler that does not have an explicit User-agent block falls back to the 'User-agent: *' section, allowing it to crawl all public pages while respecting your private path exclusions.
Open the tool:
Why blocking Google-Extended does NOT hurt your Google Search rankings
One of the most persistent misconceptions among SEO professionals is that disallowing 'Google-Extended' will damage Google organic rankings or remove the site from Google Search. This is completely false.
Google Search Central has explicitly clarified that Google-Extended is a separate product-control token. Googlebot crawls and indexes pages for Google Search, Google News, and Google Discover. Google-Extended is only used to manage whether a site's content can be utilized to train Gemini, Vertex AI generative APIs, and future model iterations.
Disallowing Google-Extended has zero impact on your core search rankings, crawling frequency, or rich snippets. If you want your site indexed by Google Search but do not want your prose used to teach Gemini how to write, blocking Google-Extended is the official, supported method.
Open the tool:
Managing aggressive crawlers: Bytespider and Common Crawl server loads
Beyond copyright and model training concerns, rogue crawlers present a major operational challenge for DevOps and server performance. ByteDance's 'Bytespider' and Common Crawl's 'CCBot' are notorious for executing tens of thousands of requests per hour on un-cached endpoints, dynamic search filters, and product paginations.
Unlike Googlebot or Bingbot, which throttle request rates when server response times rise, Bytespider has frequently been observed ignoring standard Crawl-delay directives on Apache and NGINX servers. Adding an explicit 'Disallow: /' block in robots.txt instructs legitimate spiders to cease crawling. For persistent non-compliant scrapers, you can combine this robots.txt block with Cloudflare WAF or reverse-proxy firewall rules.
Blocking CCBot additionally stops your website from being included in massive public web dumps that third-party AI companies download indiscriminately to train smaller open-source models.
Open the tool:
RFC 9309 longest-match evaluation: avoiding path conflicts
In 2022, the Internet Engineering Task Force (IETF) codified the Robots Exclusion Protocol into official internet standard RFC 9309. Prior to RFC 9309, different search engines interpreted conflicting Allow and Disallow directives unpredictably. Understanding the three core rules of RFC 9309 ensures your directives work flawlessly across all modern crawlers:
- Longest Prefix Wins: Crawlers evaluate all matching path rules and select the one with the greatest number of characters. If you specify 'Disallow: /blog/' (length 7) and 'Allow: /blog/ai-tools' (length 15), a request to /blog/ai-tools is ALLOWED because the Allow rule is longer.
- Allow Beats Disallow on Equal Length: If both an Allow and a Disallow rule match a path with identical character length (e.g. 'Disallow: /p' and 'Allow: /p'), RFC 9309 dictates that the Allow rule takes precedence.
- Wildcard (*) and End-of-Line ($) Matching: You can use asterisk (*) to match zero or more characters (e.g. 'Disallow: /*.pdf$') and dollar sign ($) to signify the end of the URL string. This allows surgical blocking of heavy file types without blocking HTML pages.
You can test these complex path matching rules instantly with the live simulator built into our AI Crawler robots.txt Generator. For optimizing agentic LLM discovery alongside robots.txt, ensure you also deploy an audited llms.txt file following our Lighthouse llms.txt guide.
Generate your custom AI crawler robots.txt in seconds
Pick your crawler strategy, toggle bots, test paths with the RFC 9309 validator, and download your ready-to-deploy robots.txt file.
Muhammad Saqlain
Cybersecurity Practitioner & Lead EngineerSecurity researcher, web developer, and founder of ToolsWebPro. Tests password entropy, GPU cracking speeds, and client-side encryption systems.
Read full author bio & credentials →Related ToolsWebPro tools
Open the free tool(s) for this guide — no signup required.
- AI robots.txt GeneratorGenerate robots.txt to block AI training scrapers like GPTBot & Bytespider while allowing ChatGPT Search and Googlebot. Free live validator in browser.
- llms.txt GeneratorGenerate and audit llms.txt & llms-full.txt files with real-time Lighthouse validation, sitemap import, and one-click fixes. Free in browser.
- YouTube OptimizerFree YouTube optimizer (TubeBuddy-style) to score titles, descriptions, and tags for better YouTube SEO, search ranking, and click-through — no signup.
More from ToolsWebPro