Paid Media

The Validity Gap

Measurement

Academic
Effort: low
Volume: high

Myth:More GRPs fix weak TV performance.

Evidence:Across 288 brands in many categories, TV ad elasticities come out far smaller than the published literature suggests, and marginal ROI is negative for more than 80% of brands.

The Nuance

The finding is about marginal dollars at observed spend levels, not that TV never works; effects vary by brand and category, and the result is robust to functional form, power, and measurement-error checks.

The Receipt

TV Advertising Effectiveness and Profitability: Generalizable Results From 288 Brands

Bradley T. Shapiro, Gunter J. Hitsch & Anna E. Tuchman, Econometrica · 2021 · Academic

Impact8.5/10
Consensus7.8/10
Evidence92/100

Channels: paid · measurement

Related Cribs

The Validity Gap

Measurement

Myth:A big enough A/B test or attribution platform can pin down each campaign's ROI.

Evidence:Across 25 large field experiments with major U.S. retailers and brokerages, the median confidence interval on ad ROI was over 100 percentage points wide; individual-level sales are so volatile (coefficient of variation ~10) that an informative test often needs 10M+ person-weeks.

Impact9.4/10
Consensus8.8/10

The Validity Gap

Budget Allocation

Myth:TV's ROI estimates justify current spend levels.

Evidence:Estimating elasticities and ROI across 288 brands, Shapiro, Hitsch & Tuchman find ad elasticities far smaller than the published literature suggests, negative ROI at the margin for more than 80% of brands, and positive overall ROI for only about a third.

Impact9.3/10
Consensus8.2/10

The Validity Gap

Attribution

Myth:Display ad ROI shows up in online conversions.

Evidence:A randomized experiment with 1.6 million Yahoo! users found display ads profitably lifted a major retailer's purchases by 5% - but 93% of the increase happened in brick-and-mortar stores, and 78% came from customers who never clicked the ads.

Impact8.8/10
Consensus8.4/10

The Validity Gap

Measurement

Myth:With enough user-level data, models can recover causal lift without experiments.

Evidence:Across 663 Facebook RCTs described by 5,000+ features, state-of-the-art observational methods (double/debiased ML, propensity matching) miss the experimental lift by a median 62-115% depending on funnel stage - larger than the median lift itself.

Impact8.8/10
Consensus8.2/10

The Validity Gap

Measurement

Myth:Valid lift measurement requires expensive PSA control ads or full platform blackouts.

Evidence:Ghost ads - logging the impression a control user would have been served, without serving or paying for it - reproduce RCT-grade measurement at a fraction of the cost of PSA controls, and are precise enough to have become standard practice at Google and for brands like Duracell and Nissan.

Impact8.7/10
Consensus8/10

Crib of the Week

One crib in your inbox every Monday. No spam, unsubscribe anytime.