ToolsWebPro logo

ToolsWebPro

AI Crawler robots.txt Generator & Validator — RFC 9309 Compliant

Create custom, RFC 9309-compliant robots.txt files to manage generative AI crawlers, model training scrapers, and search engine indexers. Block aggressive scrapers (GPTBot, Google-Extended, ClaudeBot, Bytespider, CCBot, Meta-ExternalAgent) to protect proprietary content while explicitly allowing ChatGPT Search (OAI-SearchBot), Perplexity, and Applebot for search citations and traffic. Includes preconfigured presets, custom paths, sitemap directives, and an in-browser RFC 9309 simulation tester with zero server tracking.

2026 AI Crawler Standard (RFC 9309)

Protect Your Content While Maximizing AI Search Visibility

Block unauthorized model scrapers like GPTBot, Google-Extended, and Bytespider from harvesting your content for AI training, while preserving traffic citations from ChatGPT Search (OAI-SearchBot), Perplexity, and Googlebot.

Bots Managed
13 Blocked / 10 Allowed
Customizable at any time below
OAI-SearchBotOAI-SearchBotOpenAISearch Citation

Autonomous crawler used to index web content specifically for ChatGPT Search results.

Allowing this enables your site to appear in ChatGPT Search citations and answers with direct links.
ChatGPT-UserChatGPT-UserOpenAISearch Citation

On-demand fetcher triggered when a ChatGPT user pastes your URL into a chat session.

Keep allowed so ChatGPT subscribers can summarize, cite, and read your pages in real time.
PerplexityBotPerplexityBotPerplexity AISearch Citation

Web crawler powering Perplexity AI real-time search, answer generation, and source attribution.

Critical for Perplexity SEO (AEO/GEO). Blocking eliminates high-converting citation clicks.
Claude-WebClaude-WebAnthropicSearch Citation

Real-time web browsing bot used when users ask Claude to inspect a specific URL.

Allows Claude users to read and reference your web pages on request.
ApplebotApplebotAppleSearch Citation

Apple indexer powering Siri, Spotlight suggestions, Safari lookups, and Apple Intelligence citations.

Essential for iOS / macOS search ecosystem. Disallowing removes Siri & Spotlight visibility.
GPTBotGPTBotOpenAIAI Training

Web crawler used by OpenAI to collect training datasets for future foundation models (e.g. GPT-5).

Safe to block. Blocking GPTBot does NOT impact your ChatGPT Search indexing or citations.
Google-ExtendedGoogle-ExtendedGoogleAI Training

User-agent token to opt out of Google Gemini and Vertex AI training data harvesting.

Safe to block. Does NOT affect your Google Search ranking, indexing, or Google Discover.
ClaudeBotClaudeBotAnthropicAI Training

Anthropic scraper used to train foundation models for the Claude AI family.

Safe to block. Blocking stops content scraping for model training while Claude-Web handles browsing.
Applebot-ExtendedApplebot-ExtendedAppleAI Training

Apple token allowing publishers to opt out of generative AI training for Apple Intelligence.

Safe to block. Does not block Applebot regular search index or Siri answer links.
Meta-ExternalAgentMeta-ExternalAgentMetaAI Training

Crawler used by Meta to scrape content for training Llama and other open generative AI models.

Safe to block if you want to prevent Meta AI from training on your proprietary text and code.
Meta-ExternalFetcherMeta-ExternalFetcherMetaAI Training

Secondary fetcher used by Meta AI to retrieve content for synthetic data generation.

Blocks Meta from scraping on-page resources for assistant synthesis.
cohere-aicohere-aiCohereAI Training

Training bot for Cohere enterprise LLMs and multilingual embedding models.

Safe to block. Prevents Cohere LLM dataset inclusion.
BytespiderBytespiderByteDance

ByteDance crawler for TikTok & Doubao AI models. Known for intense, high-frequency crawl rates.

Highly recommended to block. Can spike server CPU and bandwidth without providing organic traffic.
CCBotCCBotCommon Crawl

Common Crawl open dataset crawler widely downloaded by third-party AI startups and researchers.

Blocks massive open-crawl archives that distribute your text for arbitrary commercial training.
AmazonbotAmazonbotAmazon

Amazon web crawler used for Alexa answer engine and Bedrock foundation model training.

Blocks Amazon dataset mining. Set Crawl-delay or disallow if server load is a concern.
DiffbotDiffbotDiffbot

Commercial AI knowledge graph extractor that parses entities and sells structured databases.

Blocks third-party data broker harvesting your articles and products for resold APIs.
ImagesiftBotImagesiftBotImagesift

Bot specializing in harvesting images and artwork across the web for visual AI models.

Blocks non-consensual harvesting of photography, artwork, and visual brand assets.
YouBotYouBotYou.com

Crawler for You.com search and conversational agent index.

Optional block. Disallows You.com search indexing.
GooglebotGooglebotGoogle

The primary web crawler for Google organic search indexing, mobile results, and rankings.

CRITICAL: Never block this unless you want your site completely removed from Google search.
BingbotBingbotMicrosoft

Microsoft Bing crawler powering Bing, Yahoo Search, and Microsoft Copilot search indexing.

CRITICAL: Keep allowed for organic Bing and Microsoft Copilot search traffic.
DuckDuckBotDuckDuckBotDuckDuckGo

Privacy-focused search engine crawler from DuckDuckGo.

Keep allowed to maintain organic search traffic from privacy-minded DuckDuckGo users.
YandexYandexBotYandex

Primary crawler for the Yandex search engine ecosystem across Eastern Europe and Eurasia.

Keep allowed if your audience or customer base includes international or European visitors.
BaiduspiderBaiduspiderBaidu

Main search engine crawler for the Baidu ecosystem in East Asia.

Keep allowed if you target Chinese-speaking markets; block if you experience excessive bot traffic.

Standard Directives & Custom Paths

Standard sites allow all by default, then selectively disallow private routes.

Respected by Bingbot and Yandex (Google ignores crawl-delay).

robots.txt73 lines
# ============================================================================== # robots.txt generated by ToolsWebPro AI Crawler Generator (Updated: 2026-09-28) # Standard: RFC 9309 | Pure Client-Side Privacy Tool # ============================================================================== # ------------------------------------------------------------------------------ # Section 1: Disallowed AI Training Scrapers and Unauthorized Bots # ------------------------------------------------------------------------------ # GPTBot (OpenAI) - Safe to block. Blocking GPTBot does NOT impact your ChatGPT Search indexing or citations. User-agent: GPTBot Disallow: / # Google-Extended (Google) - Safe to block. Does NOT affect your Google Search ranking, indexing, or Google Discover. User-agent: Google-Extended Disallow: / # ClaudeBot (Anthropic) - Safe to block. Blocking stops content scraping for model training while Claude-Web handles browsing. User-agent: ClaudeBot Disallow: / # Applebot-Extended (Apple) - Safe to block. Does not block Applebot regular search index or Siri answer links. User-agent: Applebot-Extended Disallow: / # Meta-ExternalAgent (Meta) - Safe to block if you want to prevent Meta AI from training on your proprietary text and code. User-agent: Meta-ExternalAgent Disallow: / # Meta-ExternalFetcher (Meta) - Blocks Meta from scraping on-page resources for assistant synthesis. User-agent: Meta-ExternalFetcher Disallow: / # cohere-ai (Cohere) - Safe to block. Prevents Cohere LLM dataset inclusion. User-agent: cohere-ai Disallow: / # Bytespider (ByteDance) - Highly recommended to block. Can spike server CPU and bandwidth without providing organic traffic. User-agent: Bytespider Disallow: / # CCBot (Common Crawl) - Blocks massive open-crawl archives that distribute your text for arbitrary commercial training. User-agent: CCBot Disallow: / # Amazonbot (Amazon) - Blocks Amazon dataset mining. Set Crawl-delay or disallow if server load is a concern. User-agent: Amazonbot Disallow: / # Diffbot (Diffbot) - Blocks third-party data broker harvesting your articles and products for resold APIs. User-agent: Diffbot Disallow: / # ImagesiftBot (Imagesift) - Blocks non-consensual harvesting of photography, artwork, and visual brand assets. User-agent: ImagesiftBot Disallow: / # YouBot (You.com) - Optional block. Disallows You.com search indexing. User-agent: YouBot Disallow: / # ------------------------------------------------------------------------------ # Section 3: Default Rules for All Other Web Crawlers # ------------------------------------------------------------------------------ User-agent: * Disallow: /admin/ Disallow: /private/ Disallow: /api/ # ------------------------------------------------------------------------------ # Sitemaps # ------------------------------------------------------------------------------ Sitemap: https://yourdomain.com/sitemap.xml

Live RFC 9309 Simulator & Tester

Instant Test

Verify whether a specific crawler is allowed or blocked on any path according to your generated rules.

Access Blocked / Disallowed

Blocked by rule "disallow: /" for User-agent "gptbot".

Evaluated Rule: Disallow: /
Recommended Companion Guide

Block GPTBot but Keep ChatGPT Search (robots.txt 2026)

Read our in-depth technical analysis on the difference between OAI-SearchBot and GPTBot, how token differentiation works, and how to verify bot IP ranges to protect your server.

Read Step-by-Step Guide
✓Safety & Compliance Standards

100% Client-Side Privacy: Your domain name, sitemap paths, and generated robots.txt configuration are processed entirely in your browser using local JavaScript. No URLs, domain records, or custom rules are ever sent to ToolsWebPro servers.

How to Use AI robots.txt Generator

  1. 1Select a prebuilt crawler strategy preset (e.g. Balanced 2026 to allow ChatGPT Search while blocking AI model training, Strict Lockdown, or Permissive).
  2. 2Customize individual bot permissions (Allow, Disallow, or Default *) across AI Search bots, Foundation Model training scrapers, Aggressive crawlers, and traditional search engines.
  3. 3Configure custom standard directives including disallowed paths (e.g. /admin/, /private/), allowed paths, XML sitemaps, host, and optional crawl-delay.
  4. 4Test any target URL against your generated rules using the live in-browser RFC 9309 simulator, then copy to clipboard or download your production-ready robots.txt.

Legitimate Use Cases

Content Creators & Bloggers

Prevent foundation models from scraping your original articles for training while ensuring your articles appear with links in ChatGPT Search and Perplexity answers.

Webmasters & DevOps Engineers

Curtail aggressive bot crawls from Bytespider and commercial AI scrapers that spike server CPU, bandwidth, and hosting bills.

E-Commerce Stores & SaaS Platforms

Shield private administrative portals, checkout flows, and API endpoints from AI crawlers while feeding XML sitemaps to Googlebot.

Key Features of AI robots.txt Generator

  • Interactive Strategy Presets

    One-click configurations for Balanced 2026, Maximum AI Privacy, High-Frequency Scraper Lockdown, and Full Permissive access.

  • Granular Bot Categorization

    Differentiates between AI search citation indexers (OAI-SearchBot, PerplexityBot, Applebot) and data harvesting scrapers (GPTBot, Google-Extended, ClaudeBot, Bytespider).

  • RFC 9309 In-Browser Simulator

    Real-time simulator testing user-agents against path rules using standard longest-prefix matching and tie-breaker algorithms.

  • Server Load Protection

    Blocks bandwidth-heavy and aggressive commercial scrapers like Bytespider, Common Crawl (CCBot), and Amazonbot without harming organic SEO.

  • Search Engine Safe

    Keeps Googlebot, Bingbot, and DuckDuckBot fully allowed so your organic Google rankings and indexing remain unaffected.

  • Standard Directive Support

    Cleanly configures Disallow, Allow, Sitemap, Host, Crawl-delay, and descriptive inline comments.

  • Zero Server Uploads

    Runs 100% locally in your browser memory for complete privacy and zero data logging.

About AI robots.txt Generator

ToolsWebPro AI Crawler robots.txt Generator is a free developer and webmaster tool designed to address the modern challenge of generative AI web scraping. In 2026, web publishers must navigate two conflicting needs: preventing AI companies from scraping proprietary articles and code for model training without consent, while simultaneously allowing search-oriented AI bots (like OpenAI's OAI-SearchBot, PerplexityBot, and Applebot) to index content so users can discover their websites through conversational search. By separating model training agents (GPTBot, Google-Extended, ClaudeBot) from search citation agents (OAI-SearchBot, ChatGPT-User), this generator helps you construct valid RFC 9309 robots.txt directives that protect your intellectual property while maximizing search visibility. The built-in simulator tests your rules against any path in real time before deployment. For deeper architectural explanations, read our companion [Block GPTBot but Keep ChatGPT Search Guide](/guides/block-gptbot-keep-chatgpt-search-robots-txt-2026). If you are optimizing for agentic search engines, pair your robots.txt with our [llms.txt Generator & Validator](/tools/llms-txt-generator).

ToolsWebPro provides this utility 100% free with no account creation. Our client-side tools run locally inside your browser memory so your uploaded files are never saved or transferred to third-party servers.

Frequently Asked Questions

Is AI robots.txt Generator free on ToolsWebPro?
Yes. AI robots.txt Generator is completely free to use on ToolsWebPro. No account or payment is required.
Do I need to sign up to use AI robots.txt Generator?
No signup is needed. Open the tool, paste your link or upload your file, and get results instantly in your browser.
How does AI robots.txt Generator work?
Use AI robots.txt Generator directly in your browser without software installation or registration.
Is AI robots.txt Generator safe to use?
Yes. ToolsWebPro processes many tools client-side in your browser. We do not store your personal files on our servers for those tools.
What is the difference between GPTBot and OAI-SearchBot?
GPTBot is OpenAI's web crawler used strictly to collect training datasets for foundation models (such as GPT-4 and future iterations). In contrast, OAI-SearchBot and ChatGPT-User are used exclusively by ChatGPT Search to index content, browse live URLs on behalf of users, and generate citations and clickable source links in chat answers.
How can I block OpenAI AI training while keeping ChatGPT Search citations?
Add 'User-agent: GPTBot' followed by 'Disallow: /' to block training data scraping. Then keep 'User-agent: OAI-SearchBot' and 'User-agent: ChatGPT-User' allowed (or let them inherit your standard 'User-agent: *' allow rules). This prevents OpenAI from training on your text while preserving ChatGPT search referral traffic.
Does blocking Google-Extended hurt my Google Search ranking or SEO?
No. Google explicitly documented that Google-Extended is a standalone product token used solely to manage training data for Gemini and Vertex AI generative APIs. Disallowing Google-Extended has zero negative impact on your Google Search ranking, indexing, or Google Discover eligibility.
Why should I block Bytespider and Common Crawl (CCBot)?
Bytespider (ByteDance/TikTok) and CCBot (Common Crawl) are widely cited by webmasters as being responsible for intense, high-frequency request spikes that consume excessive server CPU and bandwidth. Furthermore, Common Crawl archives are distributed publicly and ingested by dozens of third-party AI startups without direct attribution.
What is the RFC 9309 robots.txt standard and longest prefix matching?
RFC 9309 is the official IETF standard for the Robots Exclusion Protocol published in 2022. It dictates that crawlers evaluate rules by the longest matching path prefix rather than the order in which rules appear in the file. If an Allow and Disallow rule have equal match length, Allow takes precedence.
How do I allow PerplexityBot and Applebot while blocking model scrapers?
In your robots.txt, define explicit disallow blocks for training scrapers (GPTBot, Google-Extended, ClaudeBot, Bytespider, CCBot, Meta-ExternalAgent) with 'Disallow: /'. Because PerplexityBot and Applebot do not match those user-agent blocks, they inherit your generic 'User-agent: *' permissions and remain free to index your site for citations.
Where do I place my generated robots.txt file on my server or web host?
Place the file in the top-level root directory of your website domain so it is accessible at https://yourdomain.com/robots.txt (case-sensitive, all lowercase). In Next.js, place it in public/robots.txt or generate it dynamically via app/robots.ts.
Does testing or generating robots.txt upload my domain or rules to any server?
No. ToolsWebPro executes all file generation, path parsing, and RFC 9309 simulation locally in your browser using JavaScript. No URLs, domain names, or robots.txt configurations are transmitted to our servers.

Step-by-step tutorials and troubleshooting written for this tool.

View all guides →

More free tools in the same category on ToolsWebPro.

View all hashtag generators →