43 real problems catalogued Β· New problems added weekly

Home/Tech/AI Is Answering Millions of Questions Every Day Using Content From Blogs and Publishers Who Are Getting Paid Almost Nothing For It
All problems

AI Is Answering Millions of Questions Every Day Using Content From Blogs and Publishers Who Are Getting Paid Almost Nothing For It

Anthropic's crawlers scrape a publisher's content 8,692 times for every one human visitor they send back. ChatGPT crawls websites 60,000 times and sends back one visitor. The internet's economic engine, write content, get traffic, earn advertising revenue, has been broken by AI that consumes content at industrial scale and returns almost nothing in exchange.

Added August 12, 2026
Share
8,692:1
Anthropic's scrape-to-referral ratio in Q1 2025 β€” for every 8,692 times its crawlers scraped a publisher's content, they sent back exactly one human visitor
0.33%
Click-through rate from AI chatbots to publisher websites β€” Stanford Graduate School of Business research confirming content is consumed without generating meaningful traffic or revenue
$1.5B
Amount Anthropic settled in a copyright class action in September 2025 β€” the first major financial acknowledgment that AI training on creator content without compensation has legal consequences

Problem Score

Opportunity Score

89

Strong signal β€” worth deep research.

Last verified: 2026-08-12

The Problem

The contract that nobody signed

For two decades, a simple contract held the web together. It was never written down, never formally agreed to, and operated as a shared assumption between everyone who published content online and everyone who benefited from it.

Search engines crawl your content, index it, and send humans back to your site. Those humans see your advertising, subscribe to your newsletter, buy your products. You earn enough revenue to justify producing more content. The search engine earns advertising revenue from the traffic it drives. The human gets the information they needed. Everyone benefits.

AI has broken this contract unilaterally and without negotiation.

Anthropic's crawlers scrape a publisher's content 8,692 times for every one human visitor they send back. ChatGPT crawls websites roughly 60,000 times and sends back one visitor. Perplexity's ratio is 369:1. One publisher, Digital Trends, documented 4.1 million bot scrapes in a single week that generated only 4,200 human referrals. The content is being consumed at industrial scale and returned to the publisher as a fraction of a percent of the traffic that previously funded its production.

How the economics break

The previous internet economy was built on a specific arbitrage. You produced content. Search engines indexed it. People searching found it, clicked through to your site, and generated advertising impressions. The advertising revenue covered the cost of producing more content. This is why search traffic mattered as a business metric for publishers, not because page views are inherently valuable but because page views from search generated the revenue that funded the content that generated more page views.

AI search replaces the click with a summary. A person who asks an AI system a question receives an answer synthesised from multiple sources. The AI answer is better than a list of links in many cases β€” more convenient, more direct, more immediate. For the person asking, the experience improvement is real. For the publisher whose content informed the answer, the economic consequence is real in the opposite direction. Stanford Graduate School of Business research found click-through rates from AI chatbots land at just 0.33% and AI search engines at 0.74%, compared to the rates publishers had relied upon from traditional search results.

The extraction is not proportional to any compensation. The New York Times, which has the resources and legal standing to negotiate, struck a licensing deal with OpenAI that involves payment for archive access and usage guidelines around real-time summarisation. Digital Trends, which has neither, documented 966 scrapes for every human referral with no licensing conversation ever occurring. The internet's two-decade content economy is being replaced by a new economy whose terms were set unilaterally by AI companies and whose compensation structure covers the most powerful publishers and ignores everyone else.

The legal response and its timeline problem

The litigation wave that hit the AI industry in 2025 produced results that would have seemed impossible two years earlier. Anthropic settled a copyright class action for $1.5 billion in September 2025. Reddit sued both Anthropic and Perplexity in 2025. YouTube content creators filed class actions against Nvidia, Snap, and Meta in early 2026. The New York Times case against OpenAI and Microsoft remains active.

These are genuine legal victories for content owners with the resources to pursue them. They also carry a structural problem that experts identified clearly in early 2026. Gartner estimates that 75% of AI training data in 2026 will be synthetic, potentially reaching 100% by 2030. Once AI training shifts predominantly to AI-generated synthetic data rather than human-created content, the copyright leverage that publishers are currently deploying in litigation disappears. The AI companies will no longer need human-produced content in the same way, which means the compensation claims that are currently actionable become moot.

The legal window for establishing content compensation norms is racing against the technical timeline for AI training data independence. The litigation producing settlements today may not produce the structural compensation mechanisms that would help the millions of individual creators who cannot afford to file their own cases, before AI companies stop needing the content those creators produce.

What Cloudflare changed and what it did not

On July 1, 2025, Cloudflare announced Pay Per Crawl, making it the first internet infrastructure provider to offer publishers a mechanism for charging AI crawlers for content access. The announcement was significant not only as a product but as a signal. The largest internet infrastructure company in the world had formally positioned AI crawler compensation as an infrastructure problem requiring an infrastructure solution, rather than purely a legal problem requiring litigation.

Pay Per Crawl works by requiring AI crawlers to authenticate and pay for access to content protected by the system. Publishers set their own rates. AI companies that choose to participate pay those rates. AI companies that choose not to participate can be blocked from accessing the site.

The limitation is the voluntary participation requirement. AI companies that choose not to participate can still access content through other technical means. The publishers who benefit are those using Cloudflare whose content is valuable enough to AI companies that those companies prefer paying for access to being blocked. Individual bloggers and small publishers, whose content is individually less valuable but collectively enormous, have less negotiating leverage even with Pay Per Crawl available to them.

The synthetic data cliff

The most important context for understanding the urgency of this problem is the timeline that Gartner has identified. As AI training shifts from human-created to AI-generated content, the specific compensation leverage that creators currently hold through copyright law changes character. Copyright protects original human expression. AI-generated synthetic training data is not human expression in the same legal sense.

If the transition to predominantly synthetic AI training data happens over the next three to four years, the window for establishing mandatory compensation mechanisms β€” through legislation like the AI Accountability for Publishers Act, through voluntary licensing frameworks, or through infrastructure solutions like Pay Per Crawl β€” is much shorter than the legal and legislative timelines typically require. The clock that Forbes experts described in January 2026 is running.

Proof Signals
πŸ—£οΈ
TollBit Q1 2025 State of the Bots report β€” TollBit published the most specific data available on the extraction ratio between AI scraping and human traffic referrals. OpenAI's scrape-to-referral ratio was 179:1. Perplexity's was 369:1. Anthropic's was 8,692:1. The CEO of Conductor quantified it from the publisher side: ChatGPT crawls websites roughly 60,000 times and sends back just one visitor. One publisher, Digital Trends, documented 4.1 million bot scrapes in a single week that generated only 4,200 human referrals β€” a 966:1 extraction ratio. The content is being consumed at industrial scale. The compensation is essentially zero.
πŸ—£οΈ
Cloudflare Pay Per Crawl July 2025 β€” On July 1, 2025, Cloudflare announced it was the first internet infrastructure provider to block AI crawlers accessing content without permission or compensation. The same announcement introduced Pay Per Crawl, enabling content owners to charge AI crawlers for access. The significance is not just the product but the signal: the largest internet infrastructure company in the world formally acknowledged that AI crawlers accessing content without compensation is a problem requiring a structural fix, not just a legal one.
πŸ—£οΈ
Litigation wave 2025 to 2026 β€” The New York Times sued OpenAI and Microsoft in December 2023 for copyright infringement. Anthropic settled a copyright class action for $1.5 billion in September 2025. Reddit sued both Anthropic and Perplexity AI in 2025 under multiple legal theories. YouTube content creators filed class actions against Nvidia, Snap, and Meta in early 2026. The AI Accountability for Publishers Act, introduced in February 2026, would require AI companies to get permission and pay publishers before scraping their content. The legal landscape shifted from theoretical risk to active enforcement in a single 12-month period.
πŸ—£οΈ
Forbes creator economy coverage January 2026 β€” Forbes reported in January 2026 that creators saw gains in various court claims in 2025 but experts warned the victories may be short-lived. Gartner estimates 75% of AI training data in 2026 will be synthetic, potentially hitting 100% by 2030. Once AI companies no longer need human-produced content for training, the compensation leverage creators have through litigation disappears. The window for establishing compensation norms may be shorter than the legal timeline suggests.
πŸ—£οΈ
r/blogging and independent publisher communities β€” Bloggers and independent publishers document in real time the collapse of search traffic from AI-driven search results. Posts describe sites that had stable organic search traffic for years losing 60% to 80% of that traffic within months of AI Overview rollouts, while AI systems continue to use their content to answer questions. The gap between content being consumed and content generating revenue for the creator is visible and documented at individual site level across thousands of publishers.
Who Has This Problem

The Independent Blogger

Has spent years writing detailed, researched content on a specific topic. Built a modest income from search traffic and advertising. Watches their traffic collapse as AI systems answer the questions their articles used to rank for, using their own content as source material. Has no leverage, no legal resources to file a copyright claim, and no technical infrastructure to block AI crawlers without risking blocking legitimate search engines at the same time.

The Mid-Size Publisher

Runs a publication with a small editorial team. Has licensing deals with some AI companies but the terms were negotiated under information asymmetry β€” the AI companies knew the traffic replacement risk and the publisher did not. Is now watching referral traffic decline while the licensing payment covers a fraction of the revenue the content previously generated through organic search.

The Specialist Knowledge Creator

Creates content in a specific technical or professional domain. Their content is disproportionately valuable to AI systems because it covers topics with limited available training data. Is correspondingly disproportionately scraped. Has no visibility into how their content is being used, no compensation mechanism, and no way to know whether their specific articles are being surfaced in AI answers.

The Investigative Journalist at a Small Publication

Spends weeks on original reporting that breaks a story. The story is indexed, scraped, and used to train AI systems and answer user questions about the topic. Future AI answers to questions about that topic are informed by the journalism without the journalism being credited, linked, or compensated. The economic model that funded the investigation does not scale when the traffic from the investigation is replaced by AI summaries.

Stay curious

One problem,
every Tuesday.

The most interesting problem of the week, straight to your inbox.

No spam. Unsubscribe anytime.

Why Nothing Works

Robots.txt disallow directives

Robots.txt allows website owners to instruct crawlers not to access their content. Major AI companies initially ignored robots.txt entirely. After legal and public pressure, most now claim to respect robots.txt directives. But compliance is voluntary, verification is difficult, and a publisher who blocks AI crawlers may also inadvertently affect legitimate search engine indexing. Robots.txt was designed for a world where crawling was followed by traffic referral. It was not designed for a world where crawling generates responses that eliminate the need for traffic referral.

Cloudflare Pay Per Crawl

The most credible existing infrastructure for charging AI crawlers for content access. Published in July 2025 and represents genuine progress. The limitation is that it requires AI companies to voluntarily participate in the payment system rather than simply crawl without consent. AI companies that choose not to participate can still access content through other means. Pay Per Crawl works only with willing participants, which currently means a small fraction of the AI crawlers responsible for the extraction problem.

Licensing deals with major AI companies

CondΓ© Nast, Reuters, the New York Times, and Shutterstock have struck deals with AI companies for content access. These deals compensate large publishers with the resources to negotiate and enforce licensing terms. They do not address the millions of individual creators and small publishers whose content is scraped without any licensing conversation ever occurring. The licensing ecosystem that is emerging protects the most powerful content owners and leaves everyone else without a mechanism.

Copyright litigation

Litigation has produced the most significant financial acknowledgment of the problem β€” Anthropic's $1.5 billion settlement in September 2025. But litigation requires legal resources, takes years, and produces outcomes that apply to specific plaintiffs rather than establishing automatic compensation mechanisms for all content creators. The litigation is also racing against the timeline that Gartner identifies: once AI training shifts predominantly to synthetic data, the leverage that content creators have through copyright claims diminishes significantly.

AI-generated content detection and watermarking

Tools that identify AI-generated content exist and are improving. They address the downstream problem of AI output rather than the upstream problem of AI input. A creator who can prove their work was used to train an AI model still needs a legal mechanism to be compensated for that use, which does not currently exist in most jurisdictions outside of individual settlement agreements.

Go Research This Yourself
  • πŸ”
    Security Boulevard AI content crisis analysis search: "AI scrape referral ratio TollBit publisher revenue collapse 2026"

    Published April 16, 2026. Contains the TollBit Q1 2025 scrape-to-referral ratios for OpenAI, Perplexity, and Anthropic, the Stanford GSB 0.33% click-through rate finding, and the Digital Trends 4.1 million bot scrapes documentation. The most data-dense single source available on the scale of the extraction problem.

  • πŸ”
    Forbes creator AI content claims search: "AI content creators copyright claims Gartner synthetic data 2026"

    Published January 9, 2026. Contains the Gartner 75% synthetic training data estimate, the New York Times licensing deal details, and the expert warning that the window for creator compensation claims may close as AI training shifts to synthetic data. Essential for understanding the timeline dimension of the problem.

  • πŸ”
    Tendem AI web scraping legal guide search: "AI training data legal scraping copyright Anthropic settlement 2025 2026"

    Published May 11, 2026. Contains the full litigation timeline including Anthropic's $1.5 billion settlement, the Reddit lawsuits, the YouTube creator class actions, and the AI Accountability for Publishers Act. The most comprehensive single source on the legal landscape through mid-2026.

  • πŸ”
    Cloudflare Pay Per Crawl announcement search: "Cloudflare Pay Per Crawl AI crawler compensation July 2025"

    The primary source for the July 1, 2025 Cloudflare announcement. Contains the framing that makes this a structural rather than legal problem β€” the first major internet infrastructure company treating AI crawler compensation as an infrastructure problem requiring an infrastructure solution.

  • πŸ”
    Google Trends search: "block AI crawlers, AI content theft, AI scraping compensation"

    Look at the search volume trajectory for publisher and creator responses to AI scraping since mid-2024. The growth in block AI crawlers and AI content theft searches correlates with the traffic collapse many publishers are documenting and confirms that this is an active and growing creator economy problem, not a theoretical future concern.

Questions Worth Asking
  • 1.Cloudflare's Pay Per Crawl requires AI companies to voluntarily pay for content access. What changes the incentive for AI companies to participate β€” regulatory pressure, competitive differentiation, or the risk of being locked out of the highest-quality content sources if they do not?
  • 2.Gartner estimates AI training data will be 75% synthetic in 2026 and potentially 100% synthetic by 2030. If that trajectory holds, the leverage creators have through copyright litigation and licensing negotiation disappears before any structural compensation system is established. What does the window for action actually look like?
  • 3.The licensing deals that do exist compensate major publishers. Is there a mechanism that could compensate individual creators at scale β€” something like a music streaming royalty model where a central licensing body collects payments from AI companies and distributes them proportionally to content creators based on usage?
  • 4.The AI Accountability for Publishers Act would require AI companies to get permission and pay publishers before scraping. What is the realistic path to passage given the lobbying resources available to AI companies relative to the fragmented creator and publisher community pushing for it?
  • 5.Could a browser extension or website badge that signals to AI crawlers that content is licensed rather than freely available create a practical opt-in compensation system that works without waiting for regulation or voluntary AI company participation?
⚠️ gotaprob surfaces problems worth investigating β€” not businesses ready to build. We don't validate ideas or guarantee opportunity. This is a starting point. Do your own research.

Stay curious

New problems, every week

A short digest of real problems worth exploring. No spam, no business plans β€” just the raw itch.