A customer receives a personalised recommendation, clicks it and purchases. An attribution model may assign credit to the recommendation. It still cannot tell you whether the customer purchased because of the recommendation.
The customer might have bought the same product without personalisation. The recommendation might have changed which product they selected rather than whether they bought. It might have accelerated a purchase that was coming anyway, increased a later return, reduced contribution through a discount, or affected retention months afterwards.
Attribution asks:
Which observed touchpoints should receive credit for an outcome?
Treatment-effect measurement asks:
How did the outcome differ because one treatment was selected rather than another?
Those are different questions. Attribution remains useful for journey reporting, operational optimisation and allocating observed conversions under a stated rule. It does not supply the unobserved outcome that would have occurred under a different treatment.
That missing outcome is the counterfactual at the centre of causal inference. For any customer, you observe what happened under the treatment they received. You do not simultaneously observe what the same customer would have done under another treatment. This is the fundamental problem of causal inference, and it is why effects are generally estimated across comparable groups rather than read directly from an individual journey. (1)
This guide covers the decision problem that begins after a personalisation intervention has been defined: whether treatment effects differ across eligible customers, whether targeting those differences improves policy value, and whether the policy remains profitable and safe. For the complete workflow from hypothesis and eligibility through instrumentation, implementation checks and reporting, use the complete personalisation measurement framework.
Attribution and treatment effects answer different questions
Consider a product-recommendation module.
An attribution system may record a recommendation impression, a click, a product view and an order. It may then allocate some proportion of the order to the recommendation. That allocation can answer a reporting question about the observed journey. It does not establish whether displaying the recommendation changed the probability of the order.
A treatment-effect analysis instead compares outcomes under defined alternatives:
- personalised recommendation versus a standard bestseller list;
- personalised recommendation versus no recommendation;
- algorithm A versus algorithm B;
- a discount versus a non-price benefit;
- one message time versus another;
- treatment versus suppression.
The comparison must be explicit. Personalisation is not a single treatment. A recommendation model, creative, ranking rule, discount, channel, frequency and timing can each change the intervention.
The distinction also works in the other direction. A treatment may affect an outcome without receiving any attribution. A recommendation shown today could alter a later purchase, contribution margin, product mix, return probability, repeat purchase, retention, complaints, opt-outs or customer-service contacts. An attribution window may miss those effects even when the treatment caused them.
It is worth separating this argument from browser policy, because the two are often conflated. Changes to third-party cookies affect data availability, identity linkage and outcome observation. They do not create the distinction between credit allocation and causal effect. Even with perfect tracking, an observed path still does not reveal what the same customer would have done under another treatment. Browser policy is also less settled than it is often presented: the developer of the most widely used browser stated in April 2025 that it would retain its existing third-party-cookie choice approach rather than introduce the standalone prompt it had previously planned. (2) Treat that as a changing implementation condition, not as the reason attribution differs from causal measurement.
For the wider relationship between attribution, experiments and marketing mix modelling, see the guide to measuring marketing performance. Channel and campaign lift, geo experiments and portfolio-budget decisions belong in the separate guide to marketing incrementality.
Why exposure is not causation
One of the most common personalisation comparisons is:
Customers who engaged with personalisation converted more often than customers who did not.
That comparison is usually affected by selection.
Customers who click recommendations, open messages or use personalised features may already differ in purchase intent, category interest, tenure, recent activity, available budget, loyalty, channel preference, product availability and likelihood of being reachable. The treatment may not have created those differences.
Assignment and exposure are separate events
A credible measurement system records at least:
- who was eligible;
- which treatment they were assigned;
- whether the treatment was successfully delivered;
- whether it was actually rendered or otherwise exposed;
- whether the customer complied or engaged;
- what outcome followed within the defined window.
Random assignment supports a causal comparison because treatment is not allocated according to customers’ existing propensity. Analysing customers according to assignment is commonly called an intention-to-treat comparison.
Exposure remains operationally important, but restricting the analysis to people who saw, clicked or complied with the treatment can reintroduce selection bias. Treatment itself can affect whether exposure or engagement is recorded, and conditioning on that post-assignment event can destroy the comparability the randomisation created.
Non-compliance can sometimes be handled with additional causal methods. It should not be solved by silently replacing the assigned population with an engaged subset.
Personalisation is a decision policy
Personalisation is often described as a model that predicts what a customer wants. For measurement, it is more useful to describe it as a policy.
A policy maps information available at decision time to an action: show recommendation set A, show recommendation set B, offer free delivery, offer a discount, send a reminder later, or send nothing.
The policy therefore contains more than a prediction. It determines:
- who is eligible;
- which treatments are available;
- which information may be used;
- which action is chosen;
- what constraints override the model;
- when the customer is suppressed;
- how uncertainty is handled.
One of the largest companies in the world would not commit to targeted discounting on the data it already held. We ran a custom survey to collect additional information directly from its end users, and only then did it approve a discount on a specific product group. The conversion contribution from the people who answered was strong enough to settle the argument internally, though I am not in a position to publish the figure. Which information a policy is allowed to use is a design decision rather than a technical detail. I should disclose that I now run a business built on customer-declared purchase data, so I have an interest in that argument being taken seriously.
The distinction matters because a model can predict outcomes accurately while producing a poor policy. A high purchase-propensity model may target customers who were going to purchase anyway. A high churn-risk model may target customers whose departure cannot be changed. A recommendation model may predict the item most likely to be bought while repeatedly recommending items with low incremental contribution or high return rates.
The architecture used to produce and serve recommendations belongs in the guide to AI-driven personalisation solutions. The measurement question here is narrower: whether selecting one action rather than another improves outcomes.
Customers who would have acted anyway
Suppose a customer has a high probability of purchasing under both treatment and control.
A propensity model sees an attractive target, because the predicted outcome under treatment is high. A treatment-effect model sees something different: the difference between that customer’s predicted outcome under treatment and under control may be small.
Targeting them can still be appropriate when the treatment is nearly costless and harmless. It may be wasteful when the treatment involves a discount, paid media, costly fulfilment, scarce inventory, service capacity, intrusive messaging or a risk of annoyance and opt-out.
This is the first reason personalisation attribution can overstate decision value. The attributed purchase may be entirely real. The treatment may not have created it.
Persuadable, unaffected and negatively affected groups
A useful conceptual framework considers how the same customer might behave under treatment and control.
Would act either way. The customer is likely to complete the outcome with or without treatment. Personalisation may receive attribution for the observed outcome even though the incremental effect is small or zero.
Persuadable. The treatment increases the probability of the desired outcome. This is the group a positive-uplift policy attempts to prioritise, subject to cost, uncertainty and guardrails.
Unlikely either way. The customer is unlikely to complete the outcome under either condition. Treatment adds cost or contact pressure without materially changing behaviour.
Negatively affected. Treatment reduces the probability of the desired outcome or creates another form of harm. Examples could include an irrelevant recommendation distracting from a planned purchase, a discount reducing trust or perceived quality, excessive messaging causing an opt-out, urgency framing producing a complaint, a cross-sell increasing returns, or personalisation revealing an inference the customer finds intrusive.
There is a ceiling on what any of this can do, and I spent a decade bumping into it. I watched brands with warehouse and logistics problems run well-built campaigns into late deliveries, and brands whose stock systems were fed wrong data promise availability they did not have. No recommendation model moves a customer who is unhappy for a reason the model cannot reach. Even holding the data that explained why a brand was underperforming, I had no say over its pricing, its operations or its brand. That is a commercial observation rather than a legal one, and it is why “unlikely either way” is often a statement about the business rather than about the customer.
These four patterns are conceptual descriptions of potential outcomes. They are not labels that can be observed directly for any individual person.
Why the individual counterfactual is unobservable
For a binary treatment, imagine two potential outcomes for one customer: their outcome if treated, and their outcome if not treated. Only one is ever observed.
A person who purchases after treatment could be persuadable, or could be someone who would have purchased anyway. A person who does not purchase after treatment could be unlikely under either condition, or could have been negatively affected. No historical record reveals both potential outcomes for the same customer at the same time. (1)
This limits what uplift models can claim. They estimate effects for groups of customers with similar observed characteristics, or a conditional average treatment effect under stated assumptions. They do not uncover a person’s fixed, hidden type.
Predictions for individuals can still support decisions, but they should be read as uncertain estimates derived from population patterns.
Start with a randomised holdout
Before fitting a complex uplift model, establish whether the treatment has an average effect at all.
A randomised holdout assigns eligible customers to defined alternatives independently of their expected outcome. Under a well-implemented experiment, the difference in average outcomes estimates the effect of assignment for the eligible population, and the practical literature on online controlled experiments is largely concerned with the instrumentation and process failures that break that guarantee. (3)
For a binary outcome, the basic comparison is the mean outcome in the assigned treatment group minus the mean outcome in the assigned control group. That estimate may be sufficient for many decisions.
If the average incremental contribution is positive, treatment is inexpensive and no guardrail is concerning, a broad policy may outperform a complex targeting model. If the average effect is close to zero, heterogeneity may exist, but searching for subgroups creates substantial overfitting risk.
What the holdout must preserve
The control should represent a real decision alternative, not an artificial experience nobody would otherwise receive. Possible controls include a standard non-personalised experience, the current business rule, the incumbent model, a generic recommendation, another treatment, or suppression.
The experiment must also preserve the eligible population, assignment records, stable treatment definitions, outcome availability, comparable observation windows, consistent cost accounting and guardrail measurement.
Randomisation does not correct exposure failures, missing outcomes, changing treatment versions or interference between customers.
For power, confidence intervals, p-values, sequential monitoring, multiplicity and general validity checks, use the separate guide to designing and scaling digital experiments.
Estimate the average treatment effect first
The average treatment effect answers: on average, how did assignment to this treatment change the outcome in the defined population?
It is the appropriate first result because it is directly linked to the experiment, easier to estimate than fine-grained heterogeneity, easier to explain, less vulnerable to model-selection noise, and relevant to a treat-all versus treat-none decision.
Report it with:
- the eligible population;
- treatment and comparator;
- outcome definition;
- observation window;
- absolute effect;
- uncertainty interval;
- treatment cost;
- guardrail outcomes;
- exclusions and missingness;
- treatment version.
Do not report only relative lift. A 20% relative increase from 1.0% to 1.2% is an absolute increase of 0.2 percentage points. The economic value depends on population size, contribution and cost.
For exact metric definitions and denominator discipline, use the personalisation KPI reference.
When heterogeneous treatment effects matter
An average can conceal meaningful variation. A treatment might help new customers but not established ones, increase conversion in one category and reduce it in another, work for moderate-intent customers while adding little for high-intent customers, help customers with reliable delivery options while harming those exposed to delays, or increase short-term orders among discount-sensitive customers while reducing contribution.
The conditional average treatment effect, or CATE, asks: among customers with a stated set of pre-treatment characteristics, what is the expected difference between treatment and the comparator?
The word conditional is important. It refers to averages among similar customers, not an observed individual causal effect.
Begin with justified subgroup questions
Before reaching for flexible machine learning, consider a small set of pre-specified subgroup hypotheses based on treatment mechanism, prior evidence, operational constraints, customer lifecycle, category economics, baseline risk, and known accessibility or safety considerations.
A simple subgroup estimate can be more interpretable and more stable than an unconstrained model.
Data-driven causal trees were developed partly to search for treatment-effect heterogeneity while separating discovery from estimation. Honest approaches use one part of the data to choose the splits and another to estimate the effects within them, which addresses the bias created when the most favourable-looking subgroup is both selected and estimated in the same sample. (4)
Even then, validation in later data remains necessary.
Uplift and treatment-response models
Uplift modelling estimates how the outcome changes between treatment alternatives as a function of customer characteristics. It is also described as incremental response modelling, differential response modelling, treatment-response modelling or heterogeneous treatment-effect estimation.
An uplift model should not be confused with a standard response model.
A response model estimates how likely the customer is to complete the outcome under treatment. That prediction combines baseline propensity, treatment effect, noise and model error. High predicted response does not imply high incremental response.
An uplift model estimates how much the predicted outcome differs between treatment and the comparator for customers with given characteristics. Early uplift-tree research formalised tree-based methods that split customers according to differences between treatment and control response, including settings with more than one treatment. (5)
The causal interpretation still depends on identification. Applying an uplift algorithm to confounded observational data does not make the result causal.
Model families and their assumptions
No model family is universally best. Performance depends on sample size, treatment balance, outcome prevalence, effect complexity, overlap and the quality of the underlying predictive learners.
Two-model approach, or T-learner
Fit one outcome model using treated observations and another using control observations. The estimated treatment effect is the difference between the two predictions. In meta-learner terminology this is a T-learner. (6)
It is simple, allows a flexible choice of base learner, and permits the treatment and control relationships to differ substantially. Against that, each model uses only part of the sample, errors from two models are subtracted from one another, rare treatments or imbalanced groups can produce unstable estimates, and good outcome prediction does not guarantee good effect estimation.
S-learner
Fit one outcome model containing treatment assignment as an input, potentially with treatment interactions. The effect is estimated by predicting the same customer twice, once with treatment set to each option. (6)
It uses the full sample and is operationally simple, and regularisation can stabilise the estimates. However, a flexible model may still underuse the treatment indicator, strong baseline-outcome structure can dominate a small treatment effect, and causal validity still requires valid assignment or adjustment.
X-learner
The X-learner first estimates outcomes, then constructs imputed treatment effects and models those separately before combining them. It was developed to exploit settings such as highly unequal treatment-group sizes or particular structure in the treatment-effect function. (6)
It is not automatically preferable. Its performance depends on the base learners, the weighting and the data-generating process.
Uplift trees
Uplift trees choose splits that create child groups with different estimated treatment effects, rather than merely different outcome rates. (5)
They can produce understandable segments, but small leaves are noisy, repeated split selection can overfit, treatment and control counts must remain adequate in every node, and a visually compelling tree can be unstable. Pruning and independent validation are essential.
Causal trees and causal forests
Causal trees adapt recursive partitioning to treatment effects. Causal forests average many such trees to estimate heterogeneous effects more flexibly, with asymptotic results that support confidence intervals under stated conditions. (4)(7)
Causal forests can capture non-linear interactions without specifying every subgroup in advance. Their guarantees require conditions including unconfoundedness, adequate overlap and suitable forest construction, and individual estimates can remain noisy even when the average properties are well behaved.
Generalised random forests extend the forest framework to quantities defined through local estimating equations, including heterogeneous causal parameters. (8)
Treat these as estimators, not as automatic proof that a discovered segment is real.
Doubly robust methods
Doubly robust methods combine a model for treatment assignment or logging probability with a model for outcomes under treatment alternatives.
Under the relevant assumptions, some estimators remain consistent if one of those two nuisance components is correctly specified, rather than requiring both to be correct. That is the entire content of the term. It does not mean the estimator works when both are wrong. These methods also require adequate overlap, correctly recorded treatments, appropriate features, no uncontrolled post-treatment adjustment and stable data-generating conditions.
Doubly robust construction was developed for evaluating and learning policies from logged actions and rewards, where the historical assignment probabilities are known or estimable. (9) The same doubly robust scores are used in modern policy learning, because they reduce sensitivity to any single modelling component. (10)
Complexity is not the objective
A model is not better merely because it uses more features, achieves higher predictive accuracy, produces more segments, uses a fashionable algorithm or creates a larger apparent top-decile uplift.
The relevant question is whether the resulting policy improves independently estimated business value relative to simpler alternatives.
Validate treatment-effect estimates
Treatment-effect models require a different evaluation framework from ordinary predictive models.
For outcome predictions under treatment alternatives, examine calibration, discrimination, ranking, prediction error and stability over time. These checks can reveal a poor outcome model, but they do not on their own validate a treatment effect.
For the treatment effect itself, evaluate uplift calibration, policy value, cumulative gain or Qini-style curves, later-sample performance, subgroup stability, uncertainty and sensitivity to modelling choices. Methodological work on evaluating individualised treatment-effect predictions argues that discrimination and calibration have to be defined against treatment-effect estimands specifically, and that apparent performance in development data is not an adequate substitute for independent validation. (11)
Uplift calibration
A practical calibration assessment can sort observations by predicted treatment effect, divide them into sufficiently large groups, estimate the observed treatment-control difference in each group using valid experimental or adjusted data, and compare those estimates with the average predicted effects.
This does not reveal any customer’s true effect. It assesses whether groups receiving similar predictions display compatible average effects. Small bins, rare outcomes and reused development data can all make calibration plots unstable.
Gain and Qini-style curves
A gain curve sorts customers by predicted uplift and plots cumulative estimated incremental outcomes as more of the ranked population is included. A Qini-style curve compares the model’s cumulative incremental gain with a reference such as random targeting.
These curves can be useful for assessing ranking, but they are not self-validating. The same sample should not be used freely for both model selection and final evaluation, curves require treatment-control correction, uncertainty should be shown, rankings can change across periods, and a conversion-based curve can favour an economically poor policy.
Subgroup stability
A segment is more credible when the treatment definition is unchanged, the direction and magnitude persist in later data, nearby model specifications give compatible results, the segment has enough treatment and control observations, the mechanism is plausible, and guardrails do not deteriorate.
A subgroup found once is a hypothesis until it survives validation.
Estimate policy value
A treatment-effect model produces estimates. A policy turns those estimates into actions.
A simple binary policy might be: treat the customer when expected incremental contribution is positive and sufficiently certain, otherwise use the comparator or suppress treatment.
Policy value is the expected outcome or utility achieved when decisions follow that policy. The best model by prediction error is not necessarily the best policy. Small errors near the decision threshold can matter far more than larger errors among customers who would receive the same action under either estimate.
This is why treatment choice is formalised as a welfare or value maximisation problem over a class of admissible policies, rather than as a prediction problem. (12) Policy learning from observational data uses doubly robust scores for the same reason, subject to budget, simplicity, fairness or other operational constraints. (10)
Compare with simple policies
Every learned policy should be compared with alternatives: treat nobody, treat everybody, the incumbent rule, a simple pre-specified segment, a baseline-propensity rule, a cost threshold, and a random policy at the same treatment rate.
A model that looks sophisticated but cannot beat a simple policy in later data should not be deployed because its uplift chart was persuasive.
Off-policy evaluation
Off-policy evaluation asks: using data collected under an earlier policy, what would the expected result have been under a different policy?
This can be useful before full deployment, particularly when the historical assignment included random exploration. Common approaches include a direct outcome model, inverse-probability weighting using known or estimated assignment probabilities, and a doubly robust combination of the two. (9)
Credible off-policy evaluation requires reliable logging of the historical action, a known or defensibly estimated assignment probability, sufficient overlap between historical and proposed actions, consistently observed outcomes, no unaddressed confounding, stable treatment versions, and features measured before the decision.
If the historical system almost never assigned a treatment to a particular group, its data cannot reliably evaluate a new policy that always assigns that treatment to that group. No estimator creates support where none exists.
Include treatment cost and contribution
Conversion uplift is not the same as economic value.
For a treatment and comparator, a practical decision score should consider incremental outcome probability, expected contribution if the outcome occurs, treatment cost, discount cost, fulfilment cost, service cost, return or cancellation risk, inventory constraints, future customer value and customer-harm guardrails.
Conceptually, expected incremental contribution is the incremental outcome value minus the incremental treatment and downstream cost. The exact formula depends on the use case and belongs in the financial and KPI specification rather than inside a generic model.
Why a discount can increase conversion but reduce value
Suppose a discount increases conversion by one percentage point. That does not prove the policy creates value. The additional contribution from incremental orders must exceed the discount given to customers who would have purchased anyway, the margin lost on incremental orders, fulfilment and payment costs, any increase in returns or cancellations, and possible changes in future full-price purchasing.
A propensity policy is especially vulnerable here, because it tends to give discounts to customers with high baseline intent and low incremental responsiveness.
Rank by expected incremental contribution rather than raw conversion uplift wherever reliable economic inputs exist.
For longer-horizon treatment economics, see the guide to how personalisation affects customer lifetime value.
Risk is not responsiveness
Predictive risk and causal responsiveness are distinct.
A loyal customer may be highly likely to buy under both treatment and control. Their treated outcome is easy to predict, but the treatment adds little. A customer may be very likely to leave regardless of the retention treatment, so churn risk alone does not show that intervention can retain them. A customer with moderate purchase probability may respond strongly to a relevant recommendation or service intervention. And a stable customer may be harmed by an unnecessary intervention, particularly one that is intrusive, expensive or inappropriate.
This is why ranking customers by predicted outcome under treatment is not equivalent to ranking them by predicted treatment effect, and neither is automatically equivalent to ranking them by expected incremental contribution.
Fairness and customer harm
A policy can create unequal treatment or unequal outcomes even when protected attributes are excluded from the model.
Proxy variables
Postcode, language, device, purchase history, browsing behaviour and product category can all act as proxies for protected or sensitive characteristics. Removing a protected field does not remove information correlated with it.
After a decade inside the machine, my view is that what would unsettle ordinary shoppers is not that brands record what they buy. Most people assume that. It is what gets derived from it: roughly when a salary lands, how many people someone is shopping for, which events reliably trigger a purchase. None of those is a field anyone collected. All of them are inferred, and none appears on a list of protected attributes. That is a commercial and ethical observation rather than a legal one, and it is the reason deleting a sensitive column from a model is not the same as deleting the information.
Unequal eligibility
Bias can enter before any modelling. Some customers may be excluded because their data are incomplete, unreachable through the chosen channel, absent from the experimentation population, unable to use the personalised interface, subject to different product availability, or assigned to lower-quality treatment alternatives.
Unequal model error
A model can be well calibrated overall while performing poorly for a subgroup. Treatment-effect errors may cause one group to receive unnecessary discounts, miss beneficial support, receive more intrusive contact, be exposed to inappropriate recommendations, or experience greater uncertainty around decisions that affect them.
Unequal treatment cost and harm
The same treatment may impose different burdens across customers. Frequency, timing, disclosure, financial inducements and recommendation content can all have context-dependent consequences.
Fairness constraints
Policy design may include constraints on eligibility, treatment rates, expected benefit, expected harm, model error, minimum service levels, maximum discount exposure, decision explainability and human review.
There is no single universally correct statistical definition of fairness. Different definitions can conflict, particularly where baseline outcomes differ between groups. The appropriate approach depends on the decision, the legal context, the affected population and the type of harm.
Frameworks in this area treat the question as continuing governance rather than a single model metric. The United States federal framework for artificial-intelligence risk management describes risk management across the whole system lifecycle, and is itself currently being revised. (13) The UK regulator’s guidance on artificial intelligence and data protection addresses fairness, proxies, statistical accuracy and lifecycle governance, and carries a notice that it is under review following the Data (Use and Access) Act 2025. (14) European guidance on automated individual decision-making and profiling treats these as context-dependent legal and governance questions rather than purely technical ones. (15)
For the broader data-practice context, see ethical use of consumer data in marketing.
Sensitive decisions require privacy, legal, domain and governance review. A favourable policy-value estimate does not override customer rights or safety.
Persistent controls and revalidation
A policy changes the data it later observes.
Once a model targets customers predicted to benefit, treated and untreated populations become less comparable, some actions become rare in parts of the feature space, treatment costs change, customers adapt, competitors change, product availability changes, model features drift and the treatment itself evolves.
Without continuing exploration or re-randomisation, future data may be insufficient to distinguish policy quality from selection created by the policy.
Persistent controls
A persistent control can provide an ongoing reference where withholding treatment or using the alternative is acceptable, the sample is large enough, contamination can be limited, long-term outcomes matter and treatment versions are stable.
It is not always feasible or ethical. Alternatives include periodic randomised holdouts, rotating controls, random exploration within a safe treatment set, phased launches, randomised comparisons against the incumbent policy, shadow evaluation followed by a controlled test, and stepped or capacity-based designs where appropriate.
The choice should reflect treatment risk, expected effect, sample size, customer expectations and the cost of uncertainty.
Revalidation
Revalidate when treatment content changes, eligibility changes, model features change, cost or margin changes, channel delivery changes, customer behaviour drifts, a new population is targeted, guardrails deteriorate, or the policy is expanded beyond its validated support.
A model version is not indefinitely validated because an earlier version once performed well.
Common implementation failures
- Poor randomisation. Assignment is influenced by customer characteristics, campaign operations or platform optimisation.
- Non-compliance. Assigned treatment is not delivered, rendered or used, and analysing only exposed customers creates selection bias.
- Missing outcomes. Returns, offline purchases, service contacts or delayed outcomes are absent, or differ between groups.
- Exposure errors. An impression is logged without rendering, or exposure occurs without being recorded.
- Interference. One customer’s treatment affects another’s outcome through household sharing, referrals, marketplace congestion, inventory depletion or social interaction.
- Treatment versions. Different creatives, models, discount levels or channels are grouped under one treatment label.
- Insufficient overlap. Some customer groups almost always receive one treatment, preventing defensible comparison.
- Small samples. Flexible models produce extreme subgroup estimates from few treated or control observations.
- Data leakage. Features contain information recorded after assignment, or unavailable at decision time.
- Post-treatment features. Clicks, opens and session depth are used as if they were pre-treatment characteristics.
- Multiple testing. Many outcomes, segments and model variants are searched until one looks favourable.
- Model overfitting. The policy is selected and evaluated on the same noise.
- Drift. Customers, products, delivery systems and competitive conditions change.
- Changing treatment cost. A policy trained under one discount, fulfilment or media cost is deployed under another.
- Policy changes. The historical assignment process is misrecorded or changes without versioning, undermining off-policy evaluation.
- Unstable guardrails. Conversion improves while complaints, opt-outs, returns or service contacts deteriorate.
Each of these can invalidate a sophisticated estimator. A better algorithm cannot repair an unidentified comparison or unreliable data.
Personalisation treatment-effect checklist
- Define the eligible population.
- Define the available treatments.
- Define the non-personalised or alternative policy.
- Record assignment.
- Record actual exposure.
- Define the primary outcome.
- Define treatment and downstream cost.
- Define guardrails.
- Randomise where feasible.
- Estimate the average treatment effect.
- Investigate heterogeneity using pre-specified or validated methods.
- Validate estimates on later or otherwise independent data.
- Estimate policy value.
- Compare the learned policy with simpler policies.
- Apply uncertainty, fairness and safety constraints.
- Deploy with a persistent or periodic control where appropriate.
- Monitor drift, fairness and harm.
- Revise or stop when value or safety no longer supports the policy.
Conclusion
Attributed conversions answer a credit-allocation question. They do not reveal which customers changed behaviour because of a personalised treatment.
Treatment-effect measurement begins by defining a real alternative and preserving a defensible comparison. A randomised holdout can estimate the average effect. Conditional treatment-effect and uplift models can investigate whether response varies across customers. Policy evaluation then asks whether acting on those estimates creates more value than treating everyone, treating nobody or following a simpler rule.
The individual counterfactual remains unobserved. Models do not discover a customer’s hidden type with certainty. They estimate conditional effects under assumptions, with error.
A credible personalisation policy therefore does more than rank customers by predicted conversion. It accounts for baseline propensity, incremental response, contribution, treatment cost, uncertainty, fairness, customer harm and changing conditions.
The decision is not who is most likely to act. It is who is likely to act differently because of the available treatment, whether that difference is economically and ethically worthwhile, and whether later evidence continues to support the policy.
Frequently asked questions
Is an attribution model useless for personalisation?
No. Attribution can support journey reporting, campaign operations, channel diagnostics, content reporting, platform optimisation and financial reconciliation under a stated allocation rule. What it should not be presented as is an estimate of what would have happened under a different treatment, unless the method contains a causal design that supports that interpretation.
Can an uplift model identify persuadable customers?
Not with certainty. It can estimate that customers with particular pre-treatment characteristics show a higher average treatment response under the model and its assumptions. It cannot observe both potential outcomes for one person, because only the outcome under the treatment actually received is ever recorded. Persuadable is therefore a conceptual response pattern rather than a directly observed customer label.
Do I need machine learning to measure incremental personalisation?
No. A well-designed randomised experiment and a difference in means can answer the average causal question, and pre-specified subgroup analysis may be enough for practical targeting. Machine learning becomes useful when treatment effects plausibly vary, the sample supports heterogeneity estimation, decisions must combine many characteristics, independent policy evaluation is possible, and the added complexity actually improves value.
Can treatment effects be estimated from observational data?
Sometimes, under stronger assumptions. The analysis must address why treatments differed in the historical data, using measured-confounder adjustment, instrumental variables, natural experiments or another identification strategy. Flexible machine learning does not remove unmeasured confounding; applying an uplift algorithm to confounded data produces a prediction, not a causal estimate.
Is a persistent holdout always required?
No. Persistent controls are valuable when effects, costs and customer behaviour can drift, but they may be impractical or inappropriate. Periodic experiments, rotating controls, safe exploration within a defined treatment set, or randomised comparison against the incumbent policy may be preferable. Whichever is chosen, the reasoning should be documented rather than assumed.
Does uplift modelling solve cross-channel attribution?
It avoids one particular task, which is assigning causal credit to every observed touchpoint. It does not automatically solve overlapping treatments, inconsistent assignment, missing exposure, cross-channel contamination, interference between customers, identity gaps, treatment-version changes or delayed outcomes. The treatment still has to be defined at the level of the decision being evaluated.
Do third-party-cookie changes alter the causal argument?
They affect data availability, identity linkage and outcome observation, but they do not create the distinction. Even with perfect tracking, an observed path does not reveal what the same customer would have done under another treatment. Browser policy has also proved changeable, so it is better treated as an implementation condition than as the reason attribution differs from causal measurement.
Which metric should an uplift policy optimise?
Prefer a decision-relevant measure such as expected incremental contribution, subject to safety and customer guardrails. Conversion uplift alone can favour treatments that discount customers who would have bought anyway, shift demand into low-margin products, increase returns, generate service costs, damage retention or produce unfair and intrusive outcomes.
References
-
Paul W. Holland. Statistics and Causal Inference. Peer-reviewed research article, Journal of the American Statistical Association, volume 81, issue 396, pages 945 to 960. DOI 10.1080/01621459.1986.10478354. Published December 1986. No correction, erratum or retraction notice displayed. No separate conflict declaration located on the source record; the author’s institutional affiliation was the Educational Testing Service. Not UK law. https://doi.org/10.1080/01621459.1986.10478354
-
Anthony Chavez, Google. Next steps for Privacy Sandbox and tracking protections in Chrome. Official company product-policy publication, not peer-reviewed. No reference number. Published 22 April 2025; no separate last-updated date shown. Interested-party disclosure: the publisher owns the browser and the advertising platform the announcement concerns. Product policy is subject to change and this source establishes nothing about causal measurement methodology. Not UK law. https://privacysandbox.google.com/blog/privacy-sandbox-next-steps
-
Ron Kohavi, Roger Longbotham, Dan Sommerfield and Randal M. Henne. Controlled experiments on the web: survey and practical guide. Peer-reviewed review article and practical guide, Data Mining and Knowledge Discovery, volume 18, issue 1, pages 140 to 181. DOI 10.1007/s10618-008-0114-1. Published online 30 July 2008; issue dated February 2009. No correction, erratum or retraction notice displayed. Interested-party disclosure: all four authors are listed with a Microsoft affiliation and the article draws on that company’s experimentation practice. Not UK law. https://doi.org/10.1007/s10618-008-0114-1
-
Susan Athey and Guido W. Imbens. Recursive partitioning for heterogeneous causal effects. Peer-reviewed research article, Proceedings of the National Academy of Sciences of the United States of America, volume 113, issue 27, pages 7353 to 7360. DOI 10.1073/pnas.1510489113. Published 5 July 2016. No correction, erratum or retraction notice displayed. Academic affiliations; no commercial conflict identified on the publisher record. Not UK law. https://doi.org/10.1073/pnas.1510489113
-
Piotr Rzepakowski and Szymon Jaroszewicz. Decision trees for uplift modeling with single and multiple treatments. Peer-reviewed research article, Knowledge and Information Systems, volume 32, pages 303 to 327. DOI 10.1007/s10115-011-0434-0. Published online 29 July 2011; issue dated August 2012. No correction, erratum or retraction notice displayed. Funding disclosure: supported by Polish Ministry of Science and Higher Education research grant N N516 414938; no commercial conflict statement shown. Not UK law. https://doi.org/10.1007/s10115-011-0434-0
-
Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel and Bin Yu. Metalearners for estimating heterogeneous treatment effects using machine learning. Peer-reviewed research article, Proceedings of the National Academy of Sciences of the United States of America, volume 116, issue 10, pages 4156 to 4165. DOI 10.1073/pnas.1804597116. Published 5 March 2019. No correction, erratum or retraction notice displayed. Academic affiliations; no relevant commercial relationship identified on the publisher record. Not UK law. https://doi.org/10.1073/pnas.1804597116
-
Stefan Wager and Susan Athey. Estimation and Inference of Heterogeneous Treatment Effects using Random Forests. Peer-reviewed research article, Journal of the American Statistical Association, volume 113, issue 523, pages 1228 to 1242. DOI 10.1080/01621459.2017.1319839. Published 2018. No correction, erratum or retraction notice displayed. Academic affiliations; no relevant commercial conflict identified. Not UK law. https://doi.org/10.1080/01621459.2017.1319839
-
Susan Athey, Julie Tibshirani and Stefan Wager. Generalized Random Forests. Peer-reviewed research article, The Annals of Statistics, volume 47, issue 2, pages 1148 to 1178. DOI 10.1214/18-AOS1709. Published April 2019. No correction, erratum or retraction notice displayed. Academic affiliations; the authors also develop the associated open-source implementation, which is noted here as a non-commercial interest. Not UK law. https://doi.org/10.1214/18-AOS1709
-
Miroslav Dudík, John Langford and Lihong Li. Doubly Robust Policy Evaluation and Learning. Peer-reviewed conference paper, Proceedings of the 28th International Conference on Machine Learning, pages 1097 to 1104. Published 2011. Identifiers: ACM record 10.5555/3104482.3104620; arXiv:1103.4601. No correction, erratum or retraction notice displayed. Interested-party disclosure: all authors were employed by Yahoo! Research, a party with a direct interest in internet advertising and recommendation applications. Developed for contextual-bandit data and dependent on adequate logging probabilities and support. Not UK law. https://dl.acm.org/doi/10.5555/3104482.3104620
-
Susan Athey and Stefan Wager. Policy Learning With Observational Data. Peer-reviewed research article, Econometrica, volume 89, issue 1, pages 133 to 161. DOI 10.3982/ECTA15732. Published January 2021. No correction, erratum or retraction notice displayed. Academic affiliations; no relevant commercial conflict identified on the source record. Not UK law. https://doi.org/10.3982/ECTA15732
-
Jeroen Hoogland, Orestis Efthimiou, Trang L. Nguyen and Thomas P. A. Debray. Evaluating individualized treatment effect predictions: a model-based perspective on discrimination and calibration assessment. Peer-reviewed research article, Statistics in Medicine, volume 43, issue 23, pages 4481 to 4498. DOI 10.1002/sim.10186. Published 15 October 2024. No correction, erratum or retraction notice displayed. Interested-party disclosure: one author lists an affiliation with a commercial statistical-analysis company. The work addresses randomised clinical-trial data and binary endpoints, and does not establish marketing policy value. Not UK law. https://doi.org/10.1002/sim.10186
-
Toru Kitagawa and Aleksey Tetenov. Who Should Be Treated? Empirical Welfare Maximization Methods for Treatment Choice. Peer-reviewed research article, Econometrica, volume 86, issue 2, pages 591 to 616. DOI 10.3982/ECTA13288. Published March 2018. No correction, erratum or retraction notice displayed. Academic affiliations; no relevant commercial relationship identified. Results depend on the admissible policy class and do not resolve identification from confounded data. Not UK law. https://doi.org/10.3982/ECTA13288
-
National Institute of Standards and Technology, United States Department of Commerce. Artificial Intelligence Risk Management Framework (AI RMF 1.0). Voluntary government risk-management framework, not peer-reviewed. Reference NIST AI 100-1. DOI 10.6028/NIST.AI.100-1. Published 26 January 2023. Under revision: the issuing body states that AI RMF 1.0 is being revised. Not UK law and not UK regulator guidance: it is a United States federal framework, and it is not a causal-estimation standard or a source for any particular fairness metric. https://doi.org/10.6028/NIST.AI.100-1
-
Information Commissioner’s Office. Guidance on AI and data protection. Non-statutory UK regulator guidance. No reference number. Updated 15 March 2023. Under review: the page carries a notice that, due to changes made by the Data (Use and Access) Act, the guidance is under review and may be subject to change. UK regulator guidance, not UK law, and not validation that a particular policy or estimator is fair or lawful. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/
-
Article 29 Data Protection Working Party, endorsed by the European Data Protection Board. Guidelines on Automated individual decision-making and Profiling for the purposes of Regulation 2016/679. Official regulatory guidance, not peer-reviewed. Reference WP251rev.01. Adopted 3 October 2017; last revised and adopted 6 February 2018; endorsed by the European Data Protection Board on 25 May 2018. Not UK law: this is European Union guidance, and legal interpretation depends on jurisdiction and facts. It does not define treatment-effect performance or determine whether a marketing policy is statistically fair. https://www.edpb.europa.eu/documents/guideline/automated-decision-making-and-profiling_en
Position as at 29 July 2026. Two of the governance sources above carry explicit under-review notices, browser and platform policy changes without notice, and methodological literature attracts later corrections. Check the issuing body’s current page before relying on any normative statement in this article.
Use this guide as a source
If it settled an argument in your reporting, cite it — and add Herm as a preferred source so the next answer you get from search or a model comes from work with its definitions attached.