Data-Driven Marketing

How to Build a Marketing Incrementality Testing Programme

Measure incremental marketing effects using customer holdouts, channel tests, geo experiments and lift studies, then connect results to budget decisions.

How to Build a Marketing Incrementality Testing Programme

Marketing incrementality is the causal effect of an intervention compared with what would have happened without it. The intervention may be a channel, campaign, audience rule, bid, budget level, retail-media placement or combination of activity.

That question is different from attribution.

  • Attribution allocates observed outcomes to touchpoints under a rule or model.
  • Incrementality estimates a counterfactual effect under the assumptions of a design.
  • Attributed conversions can include customers who would have converted without the advertising.
  • An incrementality estimate does not necessarily determine which touchpoint deserves narrative credit.
  • Neither method automatically identifies long-term brand effects.

Comparing large randomised advertising experiments with the observational methods normally used alongside them showed that the observational estimates could differ substantially from the experimental ones, in both directions and by amounts that varied by campaign and outcome. (1) That is the case for randomised evidence about a defined intervention and test population, although such evidence still depends on assignment, delivery, measurement and generalisation.

This article covers marketing interventions for which individual randomisation can be difficult or incomplete. For controlled changes to a website, product or defined user flow, use digital experiment design instead.

Begin with the budget or campaign decision

A useful incrementality test starts with an action, not with a platform metric.

Examples include:

  • Should paid social spend increase in these markets?
  • Does an addressable CRM channel create orders beyond those that would occur anyway?
  • What is the incremental contribution of adding retail-media activity?
  • Is a prospecting campaign incremental after accounting for customers who would already buy?
  • Does increasing search spend produce enough additional contribution to clear the marginal return threshold?
  • Should a channel remain in the portfolio, be redesigned or stop?

Write the decision in advance:

If the estimated incremental contribution for the defined population and period exceeds the required return threshold, remains acceptable under the uncertainty interval and passes the integrity checks, we will maintain or increase investment. Otherwise, we will reduce, redesign or stop it.

I sat through a lot of first meetings where a large budget was about to be signed. The stated reason was always improving the experience and lifting revenue. My habit was to ask what was actually wrong with the current setup, and the room would go quiet. They could tell something was not working, because growth and performance said so, but nobody could name it. The budget was being approved because an advanced capability was what a serious digital team was supposed to own. That is a commercial observation rather than a legal one, and it is why a test that starts from a platform metric instead of a decision usually measures the wrong thing.

Writing the decision down first also prevents a campaign from being declared successful because it generated attributed conversions while failing to change the total outcome.

Define the estimand

The estimand is the exact causal quantity the test is intended to estimate.

Possible estimands include:

  • the intention-to-treat effect of eligibility for a campaign;
  • the effect of assigning a customer to receive CRM activity;
  • the effect of making an advertising channel available in selected regions;
  • the incremental sales produced by a specified budget increase;
  • incremental contribution per eligible customer;
  • incremental return on additional ad spend;
  • the effect during the test period;
  • a cumulative effect over a pre-defined post-campaign window.

State the counterfactual condition. “No treatment” can mean no advertising, business-as-usual spend, a lower budget, another creative, a public-service announcement or a suppressed channel. These comparisons answer different questions.

Also state the population, and be honest about what the outcome data can see. An estimate among platform-identifiable eligible users does not automatically describe all customers. A geo estimate from selected regions may not transfer to a national rollout.

Something I watched for years: brands selling the same products through marketplaces with almost no information about who was buying there, while the same shopper was also buying from two competitors. Their outcome data described one shop front out of several. When the picture is that partial, programmes drift towards chasing volume with ad-hoc incentives, because there is nothing precise enough to aim at. This is a commercial point rather than a legal one, and I should disclose that I now run a business built on cross-retailer purchase data, so I have an interest in this problem being taken seriously.

Distinguish assignment, delivery and exposure

Marketing treatment is often delivered imperfectly.

A customer assigned to treatment may not receive an email. A platform may win only some auctions for treatment users. A region assigned higher spend may not achieve the planned increase. Control customers may still encounter the campaign through another device, household member, publisher or neighbouring market.

Record three layers:

  1. Assignment. The randomised condition.
  2. Delivery. Whether the planned marketing activity was delivered.
  3. Exposure. Whether the unit had an opportunity to see or experience it.

The primary intention-to-treat estimate compares outcomes by assignment. It answers the effect of the deployable policy, including normal delivery failures. An effect among actually exposed people can be useful, but exposure is often post-randomisation and platform-controlled. Estimating a treatment-on-the-treated effect requires additional assumptions or an instrumental-variable approach, and it should not replace the assignment effect without explanation.

Choose the design that matches the intervention

Terminology and platform implementations vary. The same label can conceal different eligibility rules, auction logic, control experiences, reporting windows and estimands.

  • Customer holdout. Randomises a person or account. Best suited to addressable CRM and eligible audiences. Main limitation: contamination and identity gaps.
  • Channel holdout. Randomises an eligible audience group. Best suited to the incremental effect of one channel. Main limitation: other channels may compensate.
  • Geo experiment. Randomises a region or market. Best suited to media and budget changes. Main limitation: few units, spillovers and regional differences.
  • Matched-market test. Randomises within paired regions. Best suited to activity that cannot be individually randomised. Main limitation: match quality and external shocks.
  • Conversion-lift study. Randomises within a platform-defined audience. Best suited to platform advertising. Main limitation: limited transparency and platform dependency.
  • PSA control. Randomises the ad opportunity. Best suited to separating ad exposure from targeting selection. Main limitation: cost and implementation availability.
  • Ghost-ad design. Randomises the auction opportunity. Best suited to advertising exposure effects. Main limitation: platform support and assumptions.
  • Time-based switchback. Randomises a time block. Best suited to operational or marketplace interventions. Main limitation: carryover and time trends.

No design is universally superior. Choose the one that implements the decision-relevant counterfactual with the least contamination and enough independent units to estimate it.

Customer-level holdouts

Randomly assign eligible people or accounts to treatment or control before campaign delivery. Suppress the tested activity for control while maintaining other agreed business-as-usual contact.

This design suits:

  • email, SMS, push and direct mail;
  • account-based activity;
  • loyalty or reactivation programmes;
  • addressable audiences where identity is stable.

Specify:

  • eligibility before assignment;
  • persistent assignment key;
  • household or account clustering;
  • suppression rules;
  • other channels allowed in each arm;
  • delivery and bounce handling;
  • conversion window;
  • identity matching and missing outcomes.

The main threats are contamination and identity gaps. A control customer can see the same offer through another channel or device. Members of one household or business can influence one another. Randomising at account or household level may reduce contamination but also reduces the number of independent units.

Channel holdouts

A channel holdout suppresses one channel for a randomly selected eligible group while other activity continues. It estimates the incremental effect of making that channel available under the surrounding portfolio.

That last qualification matters. Other channels may compensate. Sales teams may contact suppressed customers, paid search may capture demand that email would have converted, or an algorithm may shift budget elsewhere.

Define whether compensation is:

  • part of the policy being evaluated;
  • prohibited during the test;
  • measured as a secondary outcome;
  • a source of contamination that invalidates the contrast.

A channel can show low incremental effect because it mostly reaches customers who would convert anyway, because other channels substitute for it, or because delivery was weak. The result should lead to a mechanism review, not an automatic universal claim that the channel does not work.

Geo experiments

Geo experiments randomise non-overlapping regions to different media or budget conditions. They are useful when individual suppression is unavailable, when advertising is bought geographically, or when the business outcome is observed at regional level.

The original geo-experiment methodology assigns geographic regions to control and treatment and delivers geo-targeted advertising accordingly, estimating the effect at region level. (2) The design can produce a clear causal comparison, but the effective sample size is the number of independent regions, not the number of transactions within them.

A geo test needs:

  • stable geographic definitions;
  • sufficient independent markets;
  • a treatment that can be delivered distinctly by region;
  • baseline outcome history;
  • similarity or stratification on important pre-period features;
  • a plan for national media and shared shocks;
  • spillover assessment;
  • region-level analysis consistent with assignment;
  • a test and outcome window long enough for the intervention.

Large regions can provide many sales but few independent observations. Power depends on between-region variation, treatment intensity and the number of regions.

Matched-market tests

Matched-market designs pair regions with similar pre-period outcomes and characteristics, then assign one member of each pair to treatment.

Matching can improve precision when the match is strong, but it does not guarantee comparability after the test starts. Local promotions, weather, competitor actions, distribution changes, events or outages can affect one region.

Pre-specify:

  • the matching period;
  • variables and distance rule;
  • excluded regions;
  • assignment within pairs;
  • handling of pair loss;
  • estimator;
  • sensitivity to external shocks.

Do not choose the matching method after seeing which version produces the preferred result.

Conversion-lift studies

Platform conversion-lift studies generally form a treatment group that is shown the advertising and a control group that is held back from seeing it, within a platform-defined eligible audience, then compare downstream outcomes. They can be operationally useful where the platform controls auctions, identity and exposure.

Read the platform’s own documentation carefully rather than the summary in a reporting interface. One major platform’s help documentation states that relative conversion lift can be misinterpreted when compared across different studies, that results obtained during a study may be less accurate than end-of-study results, that very large relative lifts can arise when the control group has little conversion activity, and that modelling is used where conversions cannot be directly linked because of browser restrictions or cross-device behaviour. (3) That same page does not expressly state that assignment is random, which is worth noticing before describing the output as a randomised experiment.

The wider limitations are equally important:

  • eligibility may be platform-specific;
  • identity matching can be incomplete;
  • the platform may define the exposure opportunity;
  • auction and delivery logic may be opaque;
  • reported outcomes may exclude conversions outside the platform’s measurement system;
  • analysis choices and diagnostics may not be fully available;
  • the result concerns the tested platform implementation and period.

Treat official platform documentation as a description of an implementation, not as independent proof of effectiveness.

PSA controls

A public-service-announcement control replaces the commercial ad with a neutral or unrelated ad. The purpose is to separate the effect of the commercial message from selection into an ad opportunity.

A PSA control can improve comparability when both groups participate in similar auction and placement processes. It also consumes media inventory, can have its own behavioural effect and may be unavailable in a platform.

Document what the control ad contains, whether it changes auction cost or placement, and whether the PSA could affect the measured outcome.

Ghost-ad designs

Ghost ads use the auction opportunity that would have served an ad to construct a more comparable control. The control unit is associated with an opportunity in which the advertiser could have won, but the experimental system withholds the focal ad.

The method was developed to reduce the cost and the selection problems of measuring online ad effectiveness, in particular the expense of buying control impressions in a PSA design. (4) It requires platform or auction support and assumptions about how the counterfactual opportunity is reconstructed.

A ghost-ad estimate is not automatically portable across publishers, auction types or targeting systems.

Budget-change experiments

A budget experiment randomises a defined increase, decrease or removal of spend across regions, audiences or time periods. It is often more decision-relevant than testing advertising against none, when the real question is whether the next unit of budget is worthwhile.

Specify:

  • business-as-usual budget;
  • treatment budget or bid rule;
  • expected and achieved spend difference;
  • pacing constraints;
  • auction overlap;
  • other channel changes;
  • incremental contribution outcome;
  • the denominator for incremental return.

Incremental return on ad spend should use the incremental outcome in the numerator and the experimentally induced spend difference in the denominator. Where margin matters, use incremental contribution rather than gross revenue.

A weak treatment contrast produces weak identification. If treatment regions were assigned a 20% budget increase but delivered only 3%, report the achieved contrast and do not present the result as the effect of the intended increase.

Retail-media incrementality

Retail-media environments combine retailer identity, closed-loop sales data, sponsored placements and sometimes off-site media. Platform-reported sales can still include purchases that would have occurred without the advertising.

A retail-media test should define:

  • eligible shoppers, households or regions;
  • on-site and off-site treatment components;
  • retailer-owned and advertiser-owned outcomes;
  • identity-match coverage;
  • control suppression;
  • organic placement changes;
  • manufacturer and retailer promotions;
  • substitution across products;
  • new-to-brand or category outcomes, with exact definitions;
  • margin, fees and retailer-media cost.

Use customer holdouts where the retailer can randomise eligible accounts. Use geo, PSA, ghost-ad or budget-change designs where individual suppression is unavailable. Do not assume closed-loop reporting is causal merely because the purchase is observed.

Time-based switchback tests

A switchback alternates treatment and control across pre-defined time blocks. It can suit marketplaces, delivery systems, call centres or operational marketing interventions that affect all participants at once.

Switchbacks must address time trends and carryover. A campaign in one block may change demand in the next, weekdays differ from weekends, and supply and competitor behaviour evolve. Methodological work on switchback experiments treats the choice of block length and the resulting carryover and serial dependence as central design parameters rather than incidental details. (5)

Randomise or balance time blocks, define washout periods where needed, and analyse at the randomisation-block level.

Use the complete workflow

  1. Define the budget or campaign decision.
  2. Define the estimand.
  3. Identify the eligible population.
  4. Identify treatment-delivery constraints.
  5. Choose the randomisation unit.
  6. Assess spillovers and contamination.
  7. Estimate baseline variation.
  8. Calculate power.
  9. Pre-specify outcomes and analysis.
  10. Validate assignment and delivery.
  11. Run the test for the planned horizon.
  12. Inspect operational integrity.
  13. Estimate incremental effect and uncertainty.
  14. Translate results into contribution and cost.
  15. Determine whether the result generalises.
  16. Change, maintain or stop investment.
  17. Record the decision and revalidation need.

The steps are sequential for a reason. Power calculated before the estimand, randomisation unit and baseline variation are known is not meaningful.

Calculate power for the actual randomisation unit

Marketing experiments are often underpowered because teams count customers or impressions when treatment is assigned to a small number of markets or time blocks.

Power depends on:

  • the number of independent randomisation units;
  • allocation and stratification;
  • baseline outcome variance;
  • correlation within clusters;
  • pre-period predictive strength;
  • intended treatment intensity;
  • delivery compliance;
  • spillovers;
  • minimum decision-relevant effect;
  • outcome window;
  • multiplicity;
  • external shocks.

Advertising returns are especially difficult to estimate precisely, because incremental effects can be small relative to the variance in individual purchasing. An analysis of 25 large advertising field experiments found that obtaining precise estimates of return was statistically demanding even at very large scale, and that observational comparisons in the same settings were vulnerable to selection bias. (6) The implication is not that testing is futile. It is that the organisation must design around the uncertainty rather than promise precise answers from weak contrasts.

For geo tests, simulate or calculate power using region-level histories and the planned estimator. For customer holdouts, use the eligible customer as the unit unless assignment is clustered. For switchbacks, account for the number of time blocks, serial dependence and carryover.

Do not launch an underpowered test and then interpret a non-significant result as proof of no incrementality.

Pre-specify outcomes, windows and analysis

The primary outcome should match the investment decision. Examples include:

  • incremental completed orders;
  • incremental gross profit;
  • incremental contribution after fulfilment and returns;
  • incremental active customers;
  • incremental repeat purchase within a fixed window;
  • incremental revenue per eligible customer.

Define delayed conversions, cancellations, refunds and returns. A campaign that shifts purchases forward without increasing cumulative demand can look positive in a short window and neutral later.

Pre-specify:

  • intention-to-treat estimand;
  • assignment-level estimator;
  • covariates or pre-period adjustment;
  • uncertainty interval;
  • one-sided or two-sided test if used;
  • treatment of missing outcomes;
  • handling of failed markets or blocks;
  • multiplicity across channels, outcomes and cuts;
  • decision threshold.

The analysis should follow the randomisation. Customer-level rows do not justify customer-level standard errors when geography was randomised.

Inspect operational integrity before effect size

Check:

  • planned versus observed assignment;
  • achieved treatment-delivery difference;
  • spend and impression contrast;
  • control leakage;
  • cross-device and household contamination;
  • geo spillovers;
  • audience overlap;
  • outcome-match coverage;
  • differential missing data;
  • platform or tagging changes;
  • market shocks;
  • other campaign changes;
  • compliance with the planned horizon.

Spillovers violate the simple assumption that one unit’s treatment cannot affect another unit’s outcome. Causal inference under interference requires effects to be defined relative to a specified exposure or interference structure rather than assumed away, which is a design decision rather than something an estimator can repair afterwards. (7)

Contamination often biases the contrast towards zero, but not always. A treatment can also change competitor response, organic search, direct traffic or another channel’s algorithm. Record these mechanisms instead of treating every non-zero difference as the direct effect of the focal impressions.

Estimate incremental effect and uncertainty

A complete result includes:

  • point estimate;
  • uncertainty interval;
  • estimand;
  • population;
  • treatment and control definitions;
  • randomisation unit;
  • assignment and delivery rates;
  • time period;
  • incremental cost;
  • guardrails;
  • validity issues;
  • decision.

For example:

In eligible regions during the six-week test, assignment to the planned paid-social budget increase produced an estimated £240,000 increase in contribution relative to business-as-usual spend, with a 95% interval from minus £40,000 to £520,000. The achieved incremental spend was £160,000. The intention-to-treat incremental return estimate was 1.5, with substantial uncertainty. No pre-specified customer-service or cancellation guardrail was breached. Two regions had delivery shortfalls but remained in the assignment-based analysis. The evidence does not rule out an uneconomic effect, so the decision is to repeat with a stronger spend contrast in a second region set rather than scale nationally.

That is more useful than reporting a headline lift percentage, because it states what was changed, for whom, over what period, at what cost and with what uncertainty.

Separate short-term and long-term effects

A test period can identify only outcomes observed within its design and follow-up.

Short-run effects can differ from long-run effects because advertising may:

  • pull purchases forward;
  • create repeat customers;
  • change awareness or consideration;
  • affect price sensitivity;
  • produce category substitution;
  • alter competitor or channel behaviour;
  • decay after exposure stops.

A brand cuts price because a competitor did, and the quarter looks fine. What it has actually done is teach its customers to stop buying on their own cycle and wait for the next reaction. Sometimes you match the offer and lose the customer anyway. The check I run on a new account takes about five minutes: look at where the conversion peaks sit. If they only ever appear during sale periods, with no smaller lift around a new-season release, the customer base has already been trained. That is a commercial observation rather than a legal one, and it is precisely the kind of effect a six-week test window cannot see.

Define immediate, cumulative and post-treatment windows where the business decision requires them. Long-term holdouts can be informative but may be costly, contaminated or ethically difficult. Repeated tests and model-based extrapolation can supplement the evidence, but neither makes an unobserved long-term effect certain.

Do not label a short-term conversion-lift result as a complete brand-effect estimate.

Use incrementality and marketing-mix modelling as complementary methods

Marketing mix modelling uses aggregate time-series or panel data to estimate relationships among media, outcomes and controls under modelling assumptions. It can support portfolio allocation across channels and periods where direct randomisation is unavailable.

Experiments can provide stronger causal evidence for a defined intervention in a limited setting. Their limitations include restricted populations, short horizons, delivery constraints and low power under aggregate randomisation.

The two can work together:

  • experiment results can inform priors, calibration or validation;
  • modelling can identify channels or spend ranges where experiments would be most valuable;
  • experiments can test model-implied marginal returns;
  • disagreement can reveal different estimands, periods, populations, functional forms or data defects.

Bayesian media-mix research shows that carryover and saturation effects can be modelled explicitly, while also demonstrating that priors can strongly influence results when the sample is small and that estimated optimal allocations can carry large uncertainty. (8)

Do not average an experiment and a model estimate merely because they disagree. Reconcile the definition of treatment, outcome, population, horizon and marginal versus average effect first. Sometimes the methods are answering different questions.

For the organisation-wide method-selection framework, use marketing performance measurement.

Translate evidence into contribution and cost

A budget decision should use the economic quantity that matters.

At minimum calculate:

  • experimentally induced spend difference;
  • incremental conversions or customers;
  • incremental net revenue;
  • gross margin or contribution;
  • fulfilment, discount, fee and servicing costs;
  • incremental return with an uncertainty interval;
  • capacity or inventory constraints;
  • likely effect at the proposed scale.

Average return at the current budget is not necessarily the marginal return on the next pound. A budget-change experiment should be designed around the range in which the next decision will actually be made.

When the interval spans both harmful and attractive effects, the decision depends on reversibility, downside, opportunity cost and the value of additional information. An inconclusive result should still lead to a defined action: redesign, increase power, maintain a bounded spend, or stop because more evidence is not worth its cost.

Build a portfolio testing roadmap

Do not impose one cadence on every channel. Prioritise tests by:

  • annual spend and decision materiality;
  • uncertainty about incrementality;
  • plausibility of selection bias;
  • ability to randomise;
  • expected cost of a wrong decision;
  • treatment contrast available;
  • contamination risk;
  • time since the last valid test;
  • changes in market, platform or delivery;
  • opportunity to reuse the learning.

A practical roadmap can contain four layers.

Foundational tests. Validate identity, assignment, suppression, outcome matching and geo delivery. A small number of reliable tests is more valuable than a broad schedule built on unverified data.

Large budget questions. Test high-spend channels, major budget steps and interventions with uncertain causal value.

Portfolio interactions. Test whether channels substitute for or complement one another. A single-channel holdout estimates the effect within the surrounding portfolio, not in isolation from it.

Revalidation. Repeat material tests when targeting, platform logic, creative, pricing, competition, customer mix or measurement changes.

Maintain an experiment registry recording the decision, estimand, design, result, uncertainty, limitations, budget action and revalidation trigger. Negative, null and invalid tests belong in the repository.

Connect the programme to adjacent measurement work

Use A/B-testing validity and statistical analysis for controlled product and website experiments. The CRO diagnosis guide should generate experience hypotheses rather than own channel testing.

Personalisation measurement should own uplift and treatment-response questions for personalised experiences. Marketing analytics architecture should own data collection, identity, semantic and activation layers. First-party data measurement engineering should own consented data capture and governance.

Incrementality evidence becomes useful only when those systems connect the causal estimate to an explicit investment decision.

Frequently asked questions

What is the difference between attribution and incrementality?

Attribution allocates observed outcomes to touchpoints under a rule or model, so an attributed conversion may belong to a customer who would have purchased anyway. Incrementality estimates what would have happened without the activity, using a design whose assumptions can be stated and checked. The two answer different questions, and a channel can look strong on attributed conversions while adding little incremental contribution.

Which incrementality design should I use?

It depends on what you can randomise and what you are deciding. Use customer holdouts for addressable channels where identity is stable, geo or matched-market designs where activity is bought geographically, budget-change experiments where the real question is whether the next unit of spend pays, and switchbacks for interventions that reach everyone at once. No design is universally superior; choose the one that implements your counterfactual with the least contamination and enough independent units.

Why do incrementality tests so often come back inconclusive?

Usually because the number of independent randomisation units was small, the achieved treatment contrast was weaker than planned, or contamination pulled the comparison towards zero. Counting customers or impressions when treatment was assigned to a handful of regions overstates the effective sample. An inconclusive result is not evidence that a channel has no incremental effect; it is evidence that the study could not distinguish one.

Can I trust a platform conversion-lift result?

Treat it as a description of that platform's implementation during that period, not as independent proof. Eligibility, identity matching, exposure definition and auction logic are platform-controlled and often only partially documented, some reported conversions may be modelled rather than directly observed, and relative lift figures are not reliably comparable between studies. Read the platform's own documentation before describing the output as a randomised experiment.

How do incrementality tests and marketing mix modelling fit together?

They are complementary rather than interchangeable. Modelling can cover the whole portfolio and identify where an experiment would be most valuable, while an experiment can provide stronger causal evidence for a defined intervention and test model-implied marginal returns. When they disagree, reconcile the treatment, outcome, population, horizon and marginal versus average effect before assuming either is wrong. Do not average the two estimates.

References

  1. Brett R. Gordon, Florian Zettelmeyer, Neha Bhargava and Dan Chapsky. A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook. Peer-reviewed research article, Marketing Science, volume 38, issue 2, pages 193 to 225. DOI 10.1287/mksc.2018.1135. Published 2019. No correction, erratum or retraction notice displayed. Interested-party disclosure: two authors were employees of the advertising platform whose data and experiments were studied. Not UK law. https://doi.org/10.1287/mksc.2018.1135

  2. Jon Vaver and Jim Koehler. Measuring Ad Effectiveness Using Geo Experiments. Company research publication, 2011. No DOI and no reference number stated on the canonical page; the page does not identify the item as a peer-reviewed journal article. No publication-day or last-updated date shown beyond the year. Interested-party disclosure: issued by an advertising platform that sells the media the method measures. Not UK law. https://research.google/pubs/measuring-ad-effectiveness-using-geo-experiments/

  3. Google Ads Help. Understand your Conversion Lift based on users measurement data. Product help documentation. No reference number; help article identifier 14102450. No publication date and no last-updated date stated on the page. Publicly readable without sign-in at the date of checking. No beta, deprecation or access-restriction notice displayed; the page does record that delayed incremental conversions are currently available only for Demand Gen studies. Vendor documentation for the publisher’s own measurement product, not independent evidence of campaign effectiveness. Not UK law. https://support.google.com/google-ads/answer/14102450

  4. Garrett A. Johnson, Randall A. Lewis and Elmar I. Nubbemeyer. Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness. Peer-reviewed research article, Journal of Marketing Research, volume 54, issue 6. DOI 10.1509/jmr.15.0297. Published 2017. No correction, erratum or retraction notice displayed. Interested-party disclosure: the publisher records that two of the authors were employed by an advertising platform and held stock in it while conducting the research, and the method requires that platform’s implementation. Not UK law. https://doi.org/10.1509/jmr.15.0297

  5. Iavor Bojinov, David Simchi-Levi and Jinglong Zhao. Design and Analysis of Switchback Experiments. Peer-reviewed research article, Management Science, volume 69, issue 7, pages 3759 to 3777. DOI 10.1287/mnsc.2022.4583. Published online 1 November 2022; issue dated 2023. No correction, erratum or retraction notice displayed. Interested-party disclosure: academic authorship; disclosed research support includes an industry-affiliated research laboratory. Not UK law. https://doi.org/10.1287/mnsc.2022.4583

  6. Randall A. Lewis and Justin M. Rao. The Unfavorable Economics of Measuring the Returns to Advertising. Peer-reviewed research article, The Quarterly Journal of Economics, volume 130, issue 4, pages 1941 to 1973. DOI 10.1093/qje/qjv023. Published 6 July 2015; issue dated November 2015. No correction, erratum or retraction notice displayed. Interested-party disclosure: the publicly accessible publisher record displays no author affiliation or funding statement, and the subscriber-only published PDF and its acknowledgements were not independently verified, so no disclosure is asserted here either way. Not UK law. https://academic.oup.com/qje/article-abstract/130/4/1941/1914592

  7. Michael G. Hudgens and M. Elizabeth Halloran. Toward Causal Inference With Interference. Peer-reviewed research article, Journal of the American Statistical Association, volume 103, issue 482, pages 832 to 842. DOI 10.1198/016214508000000292. Published 2008. No correction, erratum or retraction notice displayed. Academic causal-inference research; no advertising-platform interest identified. Not UK law. https://doi.org/10.1198/016214508000000292

  8. Yuxue Jin, Yueqing Wang, Yunting Sun, David Chan and Jim Koehler. Bayesian Methods for Media Mix Modeling with Carryover and Shape Effects. Company research publication, 2017. No DOI and no reference number stated on the canonical page. No publication-day or last-updated date shown beyond the year. Interested-party disclosure: issued by an advertising and measurement provider. The work uses simulated data together with one advertiser application, and the authors show that priors can dominate small samples and that optimal media-mix estimates can carry large variance. Not UK law. https://research.google/pubs/bayesian-methods-for-media-mix-modeling-with-carryover-and-shape-effects/

Position as at 29 July 2026. Platform measurement documentation is revised without notice and often carries no publication date, methodological literature attracts later corrections, and one source above records an unverified disclosure position. Check the issuing body’s current page before relying on any specific claim in this article.

◍ herm · cite this

Use this guide as a source

If it settled an argument in your reporting, cite it — and add Herm as a preferred source so the next answer you get from search or a model comes from work with its definitions attached.

└ Erul, İ. (2026) How to Build a Marketing Incrementality Testing Programme. Herm. www.herm.io/blog/the-measurement-gap-why-92-of-marketers-miss-half-the-roi-story/
İlkem Erul
Written by

İlkem Erul

Contributor

I have over nine years of experience in digital marketing, account management, and B2C loyalty. I've helped global brands grow, and now, as a co-founder of Herm.io, I work on smarter, safer shopping experiences for consumers.

More from İlkem →

Related reading

All in this category →