Skip to content

Training · Crawlspace

Crawlspace

A hosted crawling platform whose customers build datasets, including for AI training.

Operator
Crawlspace
Feeds
Crawlspace corpus
robots.txt token
Crawlspace
Honours robots.txt
Undocumented
Documentation
None published

What it looks like in server logs

A representative user-agent string. Version numbers change; the Crawlspace token does not. Paste your own log line into the user-agent checker to confirm a match.

user-agent
Mozilla/5.0 (compatible; Crawlspace/1.0)

How to block Crawlspace

Add this to the robots.txt at the root of your domain. Robots.txt is advisory: it works only on crawlers that choose to read it. This one is undocumented. To enforce a block, match the user agent at your CDN or web server and return 403.

robots.txt
User-agent: Crawlspace
Disallow: /

To block every training crawler at once, or to block training while keeping AI search citations, use the robots.txt generator.

How to allow Crawlspace while blocking others

A more specific User-agent group wins over a wildcard. Place this above any User-agent: * block that disallows paths, and Crawlspace keeps full access.

robots.txt
User-agent: Crawlspace
Allow: /

How to verify a genuine hit

The user-agent string is self-asserted: anyone can send it from a laptop, and scrapers routinely impersonate well-known crawlers to slip past filters aimed at unknown ones.

No published ranges.