Research

Your attribution model is not measuring what you think

Paid MediaAugust 28, 20265 min read

The paper we are reading

Paywalled

A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook

Marketing Science · 2019 · 38(2), 193-225

Gordon, B. R., Zettelmeyer, F., Bhargava, N., Chapsky, D. (2019). A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook. Marketing Science, 38(2), 193-225.

Paywalled at the publisher. We have not found a free author copy.

DOI: 10.1287/mksc.2018.1135

What they found

The standard observational methods advertisers use to measure ads — matching exposed users against unexposed lookalikes, and similar statistical corrections — often failed to recover the answer the randomised experiment gave. The gap was frequently large. It was also not a consistent bias you could correct for with a fixed discount: the direction and size varied by campaign, so knowing the method overstates on average does not tell you what to do with any particular number.

How they tested it

Fifteen US advertising experiments run on Facebook, covering roughly 500 million user-experiment observations and 1.6 billion ad impressions. Each had a genuine randomised control group, giving a trustworthy answer. The authors then re-analysed the same data using the observational methods advertisers normally use, and compared what those methods produced against the experimental truth.

What it does not show

One platform, and one with unusually rich user-level data — if these methods struggle here, they will not do better elsewhere, but the specific numbers are Facebook’s. The campaigns skew toward large US advertisers. It is 2019 data, predating the privacy changes that have since made observational measurement harder still. And it tests particular observational methods; it is not a proof that every non-experimental approach fails.

Our reading

The claim, stated plainly

The number in your ads dashboard is not a measurement. It is a model output, produced by a party with an interest in the result, using a method that this study showed can miss badly against a proper control.

That is not an accusation of fraud. It is a statement about what the method can and cannot do.

Why "it overstates, so discount it" does not work

The finding that matters commercially is not that observational methods are biased upward. It is that the bias is not stable. It varied across campaigns in size and direction.

A consistent 3x overstatement would be almost fine — you would divide by three and carry on. An inconsistent one is much worse, because there is no correction available and no way to tell from the dashboard which kind of campaign you are looking at.

What we do instead

Hold out. The only reliable answer comes from a group that did not see the ads. Geo holdouts are the practical version for most budgets: withhold in a set of matched regions, run everywhere else, compare totals. Crude, cheap, and it produces a number that means something.

Watch the business number. Total revenue, total new customers, blended acquisition cost. Platform-attributed conversions are useful for steering a campaign day to day and close to worthless for deciding whether the channel deserves the budget.

Treat the sum of platform claims as a red flag. When Meta, Google and your affiliate partner each claim a share of the same purchases and the total exceeds what the business actually did, at least one of them is wrong and the dashboard will never tell you which.

Budget for measurement. A holdout costs real revenue during the test. It is still cheaper than a year of spend on a channel that was riding demand it did not create.

What this does not mean

It does not mean advertising does not work. Several of the experiments found genuine positive effects. It means the measurement most advertisers rely on is unreliable, and unreliable in a way that is not fixable by being cleverer with the same data.

The honest position is that most brands do not know their true incremental return on any given channel, and could find out for the cost of one deliberately quiet fortnight.

The study above is the work of its authors and is not ours. The summary and commentary on this page are written by Big Bang Story and are our interpretation, not the authors’. We do not host copies of other people’s papers — read it at the source.

Ready when you are

Want this applied to your spend?

We read this so it changes what we do, not so it fills a slide. Start with a free audit, or tell us what you are trying to grow.