The robots.txt Blocking Wave: What the Data Provenance Initiative Measured
When Publishers Started Saying No to Crawlers

A robots.txt file is a request rather than a lock, and the Data Provenance Initiative measured how quickly publishers began making it.
Photo: Pixabay / Pexels
A file called robots.txt sits at the root of nearly every website, a text-based instruction sheet for automated bots. A "Disallow" directive in that file tells a compliant crawler to stay away from the specified paths. The instruction carries no legal force — it is a convention, not a lock — but it is the most visible signal a publisher can send about whether it consents to having its content harvested.
Beginning in earnest in 2023, news publishers and other content owners started adding AI-specific disallow directives at a rate the internet had not seen before. The Data Provenance Initiative, a research effort examining the composition and provenance of AI training datasets, documented that rate in a 2024 study covering a large sample of web domains. Its findings showed that a substantial share of high-quality web content — the kind disproportionately represented in training corpora — became restricted within roughly twelve months, with the restriction curve steepening noticeably after OpenAI's GPTBot was announced in August 2023.

The archive lives here too now.
Photo: panumas nikhomkhai / Pexels
Among the specific crawlers being blocked: GPTBot and ChatGPT-User for OpenAI, Google-Extended for Google's generative AI products, and CCBot for Common Crawl, the open archive that has served as a foundational data source for many large language models. Publishers blocking CCBot were effectively drawing a line against a much wider range of downstream uses, because Common Crawl's datasets are incorporated by researchers and companies well beyond any single firm.
The categories of publisher that moved earliest were not evenly distributed. Large newspaper groups and digital news organisations — the outlets whose content is dense with the named facts, structured prose, and editorial judgment that AI systems most visibly exploit — were among the first movers. Academic and scientific publishers followed closely. The Data Provenance Initiative found that domains classified as high-quality by Common Crawl's own quality filters were disproportionately likely to have added restrictions, meaning the content most useful for training was also the content most rapidly being fenced off.
What robots.txt cannot do is as important as what it can. A disallow directive binds only crawlers that choose to honour it. Archival crawls conducted before the restrictions were added remain in circulation. Content already incorporated into existing training datasets is not removed by a later robots.txt update. And operators of non-compliant scrapers simply disregard the file. Publishers have understood this: robots.txt blocking has proceeded in parallel with licensing negotiations, litigation, and legislation, rather than as a substitute for any of them.


