A new research paper out of Zalando puts a number on something marketing mix modeling practitioners have long suspected but rarely quantified: standard MMM overstates paid search ROAS by roughly 2.5 times. In “Structural Estimation of Marketing Mix Model Parameters from Geo-Experiments,” posted to arXiv on August 21, Zalando’s Niklas Heusch built a structural estimation method that recovers adstock decay, saturation curvature and effectiveness directly from geo-experiment time series, differencing treatment and control regions to strip out the confounders that inflate a conventional MMM’s numbers.
The gap Heusch documents is not small. In simulation against a known ground truth of 4.20x, a realistic MMM specification returned 10.61x, badly overstating paid search’s real return. The structural approach, run across four pooled geo-tests, landed at 4.14x, close enough to the true value that its credible interval contained it. What makes the finding matter beyond one company’s methodology choice is what happened to the “oracle” version of the standard model, the one given perfect, complete confounding variables that no real practitioner ever has access to: it still returned 8.41x, nearly double the truth. Heusch’s own framing captures why: “No practitioner possesses those covariates.” A cleaner dataset does not fix the underlying identification problem.
The original insight for a marketing organization is where the fix actually lives. This is not an argument for better data hygiene inside an existing MMM, since even a hypothetically perfect one still misses by close to 2x. It is an argument that any team setting budget off a standard MMM’s paid search ROAS figure is likely over-crediting the channel by a wide margin, and the correction requires running actual geo-experiments and estimating structurally from them, not tuning the model already producing the inflated number. Teams already skeptical of the data behind their own budget calls, a skepticism this publication has covered directly among B2B marketers, now have a specific, reproducible reason: the measurement layer itself, not just the inputs feeding it, is where the number breaks. It sits alongside a broader shift toward measurement frameworks being rebuilt in public rather than trusted on faith.
Source: arXiv