Training · Open-source tool
img2dataset
The open-source tool used to download the images behind LAION and similar training sets. It does not read robots.txt but does honour X-Robots-Tag: noai response headers.
- Operator
- Open-source tool
- Purpose
- Training crawlers
- Feeds
- Dataset builder
- robots.txt token
img2dataset- Honours robots.txt
- May ignore robots.txt
- Documentation
- github.com
What it looks like in server logs
A representative user-agent string. Version numbers change; the img2dataset token does not. Paste your own log line into the user-agent checker to confirm a match.
img2datasetHow to block img2dataset
Add this to the robots.txt at the root of your domain. Robots.txt is advisory: it works only on crawlers that choose to read it. This one is may ignore robots.txt. To enforce a block, match the user agent at your CDN or web server and return 403.
User-agent: img2dataset
Disallow: /
To block every training crawler at once, or to block training while keeping AI search citations, use the robots.txt generator.
How to allow img2dataset while blocking others
A more specific User-agent group wins over a wildcard. Place this above any User-agent: * block that disallows paths, and img2dataset keeps full access.
User-agent: img2dataset
Allow: /
How to verify a genuine hit
The user-agent string is self-asserted: anyone can send it from a laptop, and scrapers routinely impersonate well-known crawlers to slip past filters aimed at unknown ones.
Not an operator; anyone can run it, from any address.
Related crawlers
Other training crawlers
Full directory of 49 AI crawlers