Do AI Companies Pay for Content They Scrape?
Almost never. Anthropic scrapes websites 73,000 times per referral sent back. Here is what publishers are doing about it and what the law says in 2026.
Almost never. AI companies scrape publisher and creator content at massive scale to train their models and answer user questions, but the overwhelming majority do so without paying anything to the people who produced that content, according to Cloudflare's crawler data published in 2026.
The scale of the imbalance is specific. Anthropic scrapes websites 73,000 times for every one referral it sends back. OpenAI scrapes 1,700 times per referral. Google, by comparison, scrapes 14 times per referral. The traditional web worked because search engines sent traffic back in exchange for access to content. AI broke that exchange without replacing it with anything.
Why AI companies scrape without paying
The legal framework that governs web scraping was not built for AI. Robots.txt, the standard way websites signal which bots can access their content, is not legally binding. AI companies can and often do ignore it.
Reddit's chief legal officer Ben Lee described the result as an industrial-scale data laundering economy. Reddit sued Anthropic in June 2025 and Perplexity in 2025 for scraping its content without a license. Perplexity claimed it did not train on Reddit data but continued citing the platform in answers. Reddit had already sent a cease-and-desist before filing suit.
The argument AI companies make is that scraping publicly available content for training falls under fair use. That argument has not been definitively tested in US courts. The cases working their way through the system in 2026 will likely determine the answer.
Stay curious
One problem,
every Tuesday.
The most interesting problem of the week, straight to your inbox.
No spam. Unsubscribe anytime.
What some publishers are getting paid
A small number of large publishers have struck licensing deals directly with AI companies:
- Reddit licenses its content to Google and OpenAI. AI licensing now accounts for close to 10% of Reddit's revenue according to Reddit's COO.
- Condé Nast, TIME, AP, Adweek, and Fortune are among the early adopters of Cloudflare's Pay Per Crawl marketplace, which lets publishers set microtransaction rates per crawl.
- The New York Times sued OpenAI and Microsoft in December 2023 for copyright infringement. The case is still active.
The pattern is consistent. Large publishers with legal resources and negotiating leverage are getting some compensation. Individual creators, bloggers, and small publishers are getting nothing.
What is changing in 2026
Two developments are shifting the landscape meaningfully.
Really Simple Licensing (RSL) is an open standard backed by Reddit, Yahoo, Medium, and Quora that builds on robots.txt to add legally enforceable licensing and royalty terms. Publishers using RSL can require AI companies to pay per crawl or per inference, meaning each time AI surfaces their content in a response. If adopted at scale, it would make ignoring content access terms a license violation rather than just a policy breach.
Cloudflare Pay Per Crawl launched in beta in 2026. New Cloudflare websites now block all AI crawlers by default. Publishers can choose to give free access, block entirely, or set a per-crawl fee. Cloudflare is considering issuing its own stablecoin to handle the microtransactions. Early adopters include some of the largest publishers in the world.
Both approaches share the same limitation. They only work if AI companies choose to participate. A company willing to ignore robots.txt can ignore these systems too.
What individual creators can do right now
The options are limited but not zero:
- Add Cloudflare to your site and configure AI crawler settings in the dashboard
- Add a clear AI training opt-out clause to your site's terms of service
- Monitor which AI crawlers are hitting your site using server logs or Cloudflare analytics
- Add your site to the robots.txt disallow list for specific AI crawlers by name
None of these guarantee compliance. They do create a paper trail that strengthens any future legal claim if an AI company scrapes your content after being told not to.
For the full breakdown of the scrape-to-referral ratios by AI company, the $1.5 billion Anthropic settlement, and why the window for establishing compensation norms may be closing faster than most creators realise, read the complete gotaprob analysis: AI Is Answering Millions of Questions Every Day Using Content From Blogs and Publishers Who Are Getting Paid Almost Nothing For It.
Sources
- Cloudflare via eMarketer — Pay Per Crawl launch, scrape-to-referral ratios by AI company including Anthropic at 73,000:1 — https://www.emarketer.com/content/cloudflare-marketplace-lets-websites-charge-ai-bots-scraping
- Reuters via TradingView — Reddit sues Perplexity for scraping data, industrial-scale data laundering quote from Reddit CLO — https://de.tradingview.com/news/reuters.com,2025:newsml_L6N3W30TM:0-reddit-sues-perplexity-for-scraping-data-to-train-ai-system
- eMarketer — Reddit, Yahoo, Medium and Quora back Really Simple Licensing standard for AI content compensation — https://www.emarketer.com/content/reddit-yahoo-others-unite-demand-ai-pay-scraped-content
Go deeper
AI Is Answering Millions of Questions Every Day Using Content From Blogs and Publishers Who Are Getting Paid Almost Nothing For It