Skip to content

Training · Diffbot

Diffbot

Feeds Diffbot's Knowledge Graph, a structured database of the web sold to enterprises and used for AI extraction. Diffbot has stated that its crawler is often operated on behalf of customers requesting specific pages.

Operator
Diffbot
Feeds
Diffbot KG
robots.txt token
Diffbot
Honours robots.txt
Undocumented
Documentation
None published

What it looks like in server logs

A representative user-agent string. Version numbers change; the Diffbot token does not. Paste your own log line into the user-agent checker to confirm a match.

user-agent
Mozilla/5.0 (compatible; Diffbot/0.1; +http://www.diffbot.com)

How to block Diffbot

Add this to the robots.txt at the root of your domain. Robots.txt is advisory: it works only on crawlers that choose to read it. This one is undocumented. To enforce a block, match the user agent at your CDN or web server and return 403.

robots.txt
User-agent: Diffbot
Disallow: /

To block every training crawler at once, or to block training while keeping AI search citations, use the robots.txt generator.

How to allow Diffbot while blocking others

A more specific User-agent group wins over a wildcard. Place this above any User-agent: * block that disallows paths, and Diffbot keeps full access.

robots.txt
User-agent: Diffbot
Allow: /

How to verify a genuine hit

The user-agent string is self-asserted: anyone can send it from a laptop, and scrapers routinely impersonate well-known crawlers to slip past filters aimed at unknown ones.

No published ranges.