Three trade publications spent the past week making sense of the same number: European news sites are absorbing far more automated AI traffic than their American counterparts. What they could not agree on, taken together, is why it is happening, and that disagreement is the more useful story for anyone budgeting content protection or weighing an AI licensing deal this quarter.

The Number Nobody Disputes

The data comes from TollBit’s State of the Bots report, which tracked 40 identified AI scraping and fetching agents across 3,906 publisher sites through the first half of 2026. Digiday framed the headline finding plainly: European sites saw roughly four times the volume of AI bot scrapes that North American sites did, and got far less back for it, one human referral for every 179 AI bot visits, against a ratio closer to 60 to 1 in North America. Scraping volume on European sites rose about 20% between January and June. Digiday’s account also carried the clearest institutional explanation on offer, quoting Grzegorz Piechota of the International News Media Association, who tied the gap to consumer demand for non-English information reshaping where AI companies source their training and retrieval data.

PPC Land’s coverage of the same report took the same dataset in a different direction, leading not with the volume gap but with a compliance gap: about 15% of identified AI page fetchers in Europe reached URLs that sites had explicitly marked disallowed in robots.txt. Three specific agents, ChatGPT-User, Bytespider and Youbot, each reached disallowed pages on nearly half the European sites that named them individually. PPC Land’s account also surfaced the sharpest tension buried in the numbers: OpenAI’s own crawler documentation argues that ChatGPT-User is not an autonomous crawler at all but a user-triggered fetch, made when a person asks a question, and that robots.txt rules written for crawlers may not govern it. TollBit’s methodology rejects that distinction outright. It counts any request that lands on a disallowed URL as a bypass, regardless of what triggered it.

Advertisement

MarTech Your brand belongs here. Reach the decision-makers who read MarTech every day. Premium placements across the site and newsletter. Advertise with us

Where the Accounts Actually Disagree

Digiday and PPC Land both treat the geographic gap itself as settled fact and argue about what sits underneath it. A third account, from ResultSense, is where the disagreement actually surfaces. It reports the same TollBit topline, European sites scraped roughly four times as often, referral ratios near 179 to 1, but then goes looking for confirmation elsewhere and does not find clean agreement. Cloudflare and DataDome, two of the largest network-layer bot-detection vendors, track much of this same traffic through their own infrastructure, and according to ResultSense their data shows inconsistent regional gaps rather than a consistent European penalty. That finding points toward publisher-level variance, site architecture, existing bot-management tooling, content-delivery setup, mattering as much as geography does. TollBit’s own explanation for the gap, offered by cofounder Olivia Joslin and cited in the ResultSense piece, is that Europe’s linguistic diversity means models training across many languages have to pull from a wider spread of European sources than a comparably sized English-language market requires. INMA’s Piechota, in Digiday, makes essentially the same argument. ResultSense is the only one of the three to put a rival dataset in the room and let it complicate the story rather than confirm it.

That is a real disagreement, not a rounding difference. If the gap is structural and linguistic, as TollBit and INMA argue, it should hold steady across any European site above a minimum size, and the fix runs through licensing and access-control decisions made at the content and legal level. If Cloudflare and DataDome are right that the regional pattern is inconsistent, the more useful unit of analysis is not “Europe” at all but individual publisher infrastructure, meaning the fix runs through the technical stack: how a site’s robots.txt is configured, what bot-management vendor sits in front of it, and how aggressively that vendor blocks by default. Those are two different capital-allocation decisions, and the three accounts collectively make clear that the industry has not settled which one applies.

A Second Data Point That Cuts Against the Language Theory

One detail inside PPC Land’s reporting sits awkwardly next to the language-diversity explanation, and none of the three accounts reconciles it. European sites disallow Claude-User at roughly a third the rate of North American sites, 9% versus 26%, and disallow Perplexity-User at roughly half the rate, 13% versus 26%. If the European gap were purely a function of AI companies needing more non-English sources to train on, blocking behavior by publishers should track scraping pressure, not run in the opposite direction. Instead, European publishers are, on average, doing less to block two of the larger AI agents than North American publishers are, which is at least as consistent with the Cloudflare and DataDome view, that publisher-side configuration choices are doing more work than geography, as it is with the TollBit and INMA explanation. None of the three outlets frames it that way directly. Reading them side by side is what surfaces it.

Newsletter

Get the week's best tech coverage.

Free. Read by thousands of HR, tech, and business leaders.

The Compliance Gap Nobody Can Enforce Alone

The second thread running through all three accounts is arguably the bigger near-term problem: robots.txt, as currently implemented, functions as a request rather than an enforcement mechanism. A publisher can write a disallow rule for every AI agent it can name, and PPC Land’s reporting shows that still leaves a meaningful share of fetches landing on pages marked off-limits, either because an operator disputes that the rule applies to it, because a new agent has not been named yet, or because compliance is simply inconsistent in practice. TollBit’s March data, cited across the coverage, found the share of bots ignoring disallow directives on its network jumped from 3.3% to 12.9% in a single quarter. That is not a rounding error either, and it lands alongside a separate track of US legislative activity aimed at forcing AI crawlers to identify themselves honestly, which suggests lawmakers on at least one continent view voluntary disclosure as already broken. It is a signal that voluntary compliance is degrading precisely as more AI products route through fetch-on-demand agents rather than batch crawlers, which is a harder pattern to block without also blocking the referral traffic publishers still want.

What It Means for the Marketing Leader

None of this is abstract for a marketing or content organization running a European site, or a US-based brand syndicating content into European markets. First, robots.txt should be treated as documentation of intent, not a technical control; anyone relying on it as the sole layer of content protection is already exposed, and the compliance gap the coverage describes will not close on its own. Second, before signing off on new bot-management spend, ask which of the two explanations your own traffic data supports: a consistent European penalty regardless of setup, which argues for licensing and legal-side deals with AI platforms, or an inconsistent, configuration-driven gap, which argues for tightening the technical stack first. Guessing wrong means spending on the wrong lever. Third, the OpenAI dispute over what counts as a “crawler” versus a “user-initiated fetch” is not a technicality. It is the industry’s live argument over whether robots.txt governs AI agents at all, and marketing and legal teams negotiating AI licensing terms should assume that argument is unresolved rather than assume the old rules of engagement still apply. Fourth, do not let the specific bot names in this quarter’s data set a false sense of completeness in a content-protection policy. ChatGPT-User, Bytespider and Youbot are this report’s worst offenders on disallowed-URL access, but the compliance-rate trend TollBit describes, bots ignoring disallow directives roughly quadrupling in a single quarter, says the roster of offenders is moving faster than any static blocklist can track. A policy built around today’s named agents is stale by the next quarterly report.

Publishers are not standing still on the monetization side either; some are already restructuring how they license access to agentic systems rather than simply trying to block them, which is the practical middle path between the licensing-heavy and infrastructure-heavy responses this week’s coverage implies. The honest summary of this week’s coverage is not that AI scraping in Europe is a settled crisis with one clear cause. It is that three credible outlets read the same report and produced three different diagnoses, and the gap between those diagnoses is exactly where a publisher’s next budget decision should start.

Source: TollBit