Skip to content

Training · ByteDance

Bytespider

ByteDance's crawler, associated with the Doubao model and TikTok. It is among the highest-volume AI crawlers on the web, and many site operators report it continuing to fetch pages after being disallowed in robots.txt.

Operator
ByteDance
Feeds
Doubao / TikTok
robots.txt token
BytespiderAlso: TikTokSpider
Honours robots.txt
Undocumented
Documentation
None published

What it looks like in server logs

A representative user-agent string. Version numbers change; the Bytespider token does not. Paste your own log line into the user-agent checker to confirm a match.

user-agent
Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; [email protected])

How to block Bytespider

Add this to the robots.txt at the root of your domain. Robots.txt is advisory: it works only on crawlers that choose to read it. This one is undocumented. To enforce a block, match the user agent at your CDN or web server and return 403.

robots.txt
User-agent: Bytespider
User-agent: TikTokSpider
Disallow: /

To block every training crawler at once, or to block training while keeping AI search citations, use the robots.txt generator.

How to allow Bytespider while blocking others

A more specific User-agent group wins over a wildcard. Place this above any User-agent: * block that disallows paths, and Bytespider keeps full access.

robots.txt
User-agent: Bytespider
Allow: /

How to verify a genuine hit

The user-agent string is self-asserted: anyone can send it from a laptop, and scrapers routinely impersonate well-known crawlers to slip past filters aimed at unknown ones.

ByteDance publishes no IP list.