# AI crawler rules: block training only # Generated from the NumberHill AI crawler directory (24 agents, list updated 2026-09-09) # https://numberhill.com/tools/ai-robots-txt # Paste above any "User-agent: *" group. Robots.txt is advisory; see each crawler's page for whether it is honoured. User-agent: GPTBot User-agent: ClaudeBot User-agent: anthropic-ai User-agent: Claude-Web User-agent: Google-Extended User-agent: GoogleOther User-agent: GoogleOther-Image User-agent: GoogleOther-Video User-agent: Applebot-Extended User-agent: Amazonbot User-agent: Meta-ExternalAgent User-agent: FacebookBot User-agent: CCBot User-agent: Bytespider User-agent: TikTokSpider User-agent: cohere-ai User-agent: Diffbot User-agent: PanguBot User-agent: ImagesiftBot User-agent: AI2Bot User-agent: AI2Bot-Dolma User-agent: omgili User-agent: omgilibot User-agent: Webzio-Extended User-agent: YandexAdditional User-agent: YandexAdditionalBot User-agent: SBIntuitionsBot User-agent: ICC-Crawler User-agent: VelenPublicWebCrawler User-agent: img2dataset User-agent: aiHitBot User-agent: Crawlspace User-agent: Poseidon Research Crawler Disallow: /