Measuring Performance

How to Measure the Effectiveness of Personalisation

Build a personalisation measurement framework using exposure definitions, control groups, incrementality and instrumentation checks.

How to Measure the Effectiveness of Personalisation

Most personalisation programmes are not badly built. They are badly compared.

I spent close to a decade on the account side of enterprise personalisation, in Istanbul and then Paris, and when a client told me their programme was underdelivering I got into the habit of guessing the cause before opening the data. It was almost always one of two things: the analytics were being read wrongly, or the test had been run without a proper setup. Genuinely weak models came a distant third.

This guide is about the second problem. It covers how to design and operate a personalisation measurement framework: what to define before you start, how to build a comparison you can defend, how to tell a causal claim from a descriptive one, what has to be true about your instrumentation before any number means anything, and how findings turn into decisions.

It does not define individual metrics. For formulas, denominators, required data and interpretation warnings on 19 specific KPIs, see our personalisation KPIs and metrics reference.

What personalisation effectiveness actually means

Effectiveness is not “did the personalised experience perform well”. It is “did the personalised experience produce an outcome that would not have happened otherwise, at a cost worth paying”.

Those are different questions, and the gap between them is where most personalisation reporting lives.

Consider what the widely quoted industry figures actually measure. McKinsey’s 2021 research is cited constantly as evidence that personalisation leaders generate 40% more revenue. Read the methodology exhibit and the picture changes: the sample was 20 consumer companies without direct customer relationships, respondents were split into faster- and slower-growing halves, and the question asked what percentage of their revenue came from personalised marketing actions (1)(2). The measured variable was the self-reported share of revenue attributed to personalisation, not total revenue, and the comparison was between growth rates rather than between a treatment and a control.

The same paper’s consumer figures hold up better. In a September 2021 survey of 1,013 US adults, 71% at least somewhat agreed that they expect personalised interactions and 76% at least somewhat agreed that irrelevant recommendations are frustrating (1). That is a US attitude survey, measured on a six-point agreement scale, and it tells you what people say they expect. It does not tell you what personalisation causes.

None of this means personalisation does not work. It means the evidence most often used to prove it works does not, in fact, prove it, and a measurement framework that starts from those numbers has started from the wrong place.

How many organisations can actually measure the return? There is no reliable overall figure. The closest current source is a 2025 Forrester study commissioned by Adobe, covering 647 personalisation decision-makers, in which 14% of the organisations classified as most mature and 27% of the least mature selected “we are unable to assess ROI” as a challenge (3). No overall percentage is published, and not selecting that option is not evidence that an organisation measures ROI well.

Start from the decision, not the metric

Before anything else, write down the decision the measurement exists to support. Not the metric. The decision.

There are only a few decisions personalisation measurement ever supports:

  • Should we deploy this experience more widely, keep it as it is, or stop it?
  • Which of two or more approaches should we use?
  • Is this programme worth what it costs?
  • Is something broken?

Each demands a different comparison. “Should we keep sending this?” compares the campaign against no contact. “Does personalisation beat our standard experience?” compares the personalised journey against the best credible non-personalised alternative. “Which ranking strategy wins?” holds channel, creative and frequency constant and varies only the rule.

Teams that skip this step end up with a dashboard that answers none of the three, because the same number was asked to serve all of them.

Define eligibility, exposure and outcome before you start

Three definitions have to exist in writing before a single number is collected.

Eligibility is who could have received the personalised experience. It is determined by consent, identity resolution, channel availability and whatever business rules apply. It has to be recorded before treatment, because it is the population your result generalises to.

Exposure is who actually received it. A decision returned by a server is not an exposure. The page may have closed, the module may not have rendered, the content may have loaded below the fold.

Outcome is the behaviour that has to change for this to have worked, with an observation window long enough for that behaviour to appear.

The distinction between eligibility and exposure is the single most consequential thing in this article, and it is where the most common error in the field lives.

Why comparing exposed and unexposed customers does not work

It is intuitive to compare people who saw the personalisation with people who did not. It is also close to useless, because those two groups differ for reasons that independently predict the outcome.

Customers are unexposed because they lacked consent, were unidentified, arrived through a different channel, or were less interested to begin with. Customers who clicked a recommendation had already demonstrated intent. Conditioning on exposure selects people partly on the basis of the very thing you are trying to measure.

The best available estimate of how large this problem gets comes from an instrumental-variable study of Amazon browsing covering 2.1 million users over nine months. The researchers concluded that at least 75% of recommendation click-through activity would probably have occurred without the recommendations at all (4).

That figure is specific to its context and should not be applied as a correction factor to anyone else’s programme. What travels is the direction and the scale. Attributed recommendation traffic and incremental recommendation traffic are different quantities, and the gap between them can be most of the number.

The version of this I watched happen

A French beauty retail chain I worked with ran what looked like a textbook loyalty automation. Thirty days after a purchase, customers in the loyal segment received a discount. For five months the reporting was excellent: strong opens, strong clicks, healthy conversion. Nobody had a reason to question it.

Then we split the audience and held a group back with no discount at all. Opens came out similar. Clicks were actually lower in the discount group. And the difference in conversion between receiving the discount and receiving nothing was about two percentage points.

Five months of good-looking reporting, and what the programme had mostly been doing was handing money to customers who were going to buy anyway.

Nothing in the dashboard could have revealed that, because the dashboard contained no counterfactual. Only the holdout did.

I am describing client work under confidentiality, so the brand is not named, and the figures are from my own recollection rather than a published result. The mechanism is the point, not the number.

Build the control group

A control group is a randomly assigned set of eligible customers who are deliberately excluded, so their behaviour can stand in for what would have happened anyway.

Randomise at customer or account level, not message level. Message-level assignment lets the same person appear in both conditions through different channels.

Assign before or at the moment of eligibility. Never let a customer’s later response determine which group they land in. If the assignment happens after someone engages, you have rebuilt the selection problem inside your experiment.

Make the holdout persistent across channels. A control customer who is suppressed from email but still sees the personalised web module is not a control.

Analyse by assignment, not by exposure. Report the intention-to-treat difference between assigned groups. Dropping assigned customers who never loaded the experience reintroduces exactly the bias the randomisation removed. Exposure rate is a useful implementation diagnostic; it is not a safe basis for the causal comparison.

Where randomisation is genuinely impossible, a phased rollout with a credible untreated comparison group is the next best thing. Describe it as observational, because it is.

Establish a baseline that means something

A baseline is not last month’s number. It is the same metric, computed the same way, on the same population, over a period long enough to contain the normal variation of your business.

Two things break baselines routinely.

The first is a definitional change nobody logged. If the engagement definition, the attribution window or the eligibility rule moved during the measurement period, the comparison is invalid regardless of how the experiment was run.

The second is that a strong before-and-after is not an experiment. At one of the biggest menswear brands in Türkiye, campaigns had been reporting conversion uplift in the 7 to 12% range. After we introduced structured A/B testing and gamified mechanics, uplift reached around 25%. It remains the largest before-and-after I personally saw in ten years, and I was proud of it.

It is still a before-and-after. Season changed, the product mix changed, the site changed, and the team had got better at the work. The A/B testing we introduced is what made subsequent results defensible. The headline number itself is an observation about one brand in one category, and I would not offer it to anyone as a benchmark.

Measure incrementality, and say so only when you have

The word “incremental” carries a causal claim. Reserve it.

Use “difference”, “association” or “observed lift” for descriptive comparisons. Use “incremental”, “caused by” and “attributable to” only for results backed by random assignment or another defensible causal design.

This is not pedantry. Once a descriptive number is labelled incremental in a board pack, it gets multiplied by the eligible population and becomes a business case.

Two further disciplines matter.

Report absolute effects alongside relative ones. A rise from 5% to 6% is one percentage point and a 20% relative lift. Both are true and they communicate very different magnitudes.

Correct for selection when summing wins. If you ship only the experiments that came back significantly positive and then add their estimates together, the total will overstate the programme. Airbnb documented this as a winner’s curse problem and described using a persistent split holdout as a meta-experiment, comparing the combined measured effect against the sum of individual winning estimates, with a mathematical correction for the selection-induced bias (5). If you run many personalisation tests, your annual programme number needs one of those two safeguards.

Choose between experiments and observational analysis

Experiments are better for causal questions. They are not always available and they are not always adequate.

Use a randomised experiment when the decision is causal, the outcome occurs often enough to detect, and you can hold a control group without unacceptable commercial cost.

Use observational analysis for describing what happened, generating hypotheses, diagnosing where a journey breaks, and sizing populations. State plainly that it is descriptive.

Be honest about the limits of experiments too. Effects on long-term outcomes are frequently too small to detect at ordinary sample sizes, which is a statistical fact rather than evidence of no effect.

For attribution modelling, marketing-mix modelling and geo-experiments, which sit outside personalisation specifically, see our guide to measuring marketing performance.

Choose a metric hierarchy, not a metric

No single number describes a personalisation programme, and no single statistic is the misleading one. Looking at any statistic in isolation is what misleads.

I sat through a great many reviews where a conversion-rate uplift was presented as the result and nobody had checked what happened to average order value underneath it. Occasionally the answer was that order value had moved the other way and the revenue effect was negative. The conversion number was accurate. It was also the wrong thing to have looked at alone.

Build three layers:

Delivery. Coverage, exposure, decision latency, failure rate. These explain why an aggregate effect is smaller than an experimental result. A model can only influence the opportunities where it actually runs.

Behaviour and commerce. Conversion, revenue per visitor, order value, contribution margin. This is where the decision usually gets made.

Guardrails. Opt-outs, complaints, returns, incentive cost, repeat service contacts, product diversity.

Then decide, before the test starts, which single metric is primary. The rest are diagnostics.

The long-term outcome problem

The outcome businesses actually care about is retention, and retention is the hardest thing to move detectably with a recommendation change.

Netflix has described this openly. It reports that retention is affected by external factors like seasonality and personal circumstances, is mainly sensitive among customers already close to cancelling, may follow a sequence of poor experiences rather than one, and produces only one signal per account per month. Its conclusion is that optimising for retention alone is impractical, so it trains against more sensitive proxy rewards while treating retention as the north star (6).

The trap in that approach is validating the proxy the wrong way. Netflix’s own subsequent research is explicit: a user-level correlation between engagement and retention is not evidence that increasing engagement will increase retention, and the correlation between estimated treatment effects can itself be biased when both metrics carry correlated sampling error. In their application across 96 treatment-control comparisons, downsampling experiments from 15 million units to one million made the naive relationship between short- and long-term metrics roughly 50% steeper (7).

The practical rule: validate a proxy against treatment effects across your previous experiments, not against the fact that engaged customers retain longer.

Segments, journeys and the limits of a result

A result belongs to the population it was measured on. Two things follow.

Analyse subgroups only on characteristics defined before treatment. Splitting results by post-treatment behaviour, such as whether someone engaged, recreates selection bias inside a valid experiment.

Do not assume a tactic transfers across categories. In fashion, easing the path to the basket and sending ad traffic straight to category pages worked reliably for us. We tried the same treatment for a French automotive brand and bounce rates went up slightly. We stopped the campaign and spent the effort on making product detail pages easier to read instead, which improved both bounce and overall engagement.

Fashion buyers want to get to the basket. Car buyers want to read and compare. The tactic was not wrong; the transfer was. A personalisation result from a low-consideration category is a hypothesis in a high-consideration one, not a finding.

Validate the instrumentation before you interpret anything

There is a category of measurement failure that has nothing to do with statistics, and it invalidates results silently.

Sample ratio mismatch occurs when the observed split between variants differs from the configured split. It is a symptom rather than a diagnosis, and its documented causes run through the whole experiment lifecycle: incorrect bucketing, unstable identifiers, variant-specific redirects, treatment-induced changes in whether users trigger or report, dropped events, incorrect joins, conditioning on treatment-affected variables, and uneven ramping (8).

Microsoft treats this as a hard gate. Its platform runs a chi-square test on user counts before revealing any experiment outcome, at a conservative threshold, and it reports that roughly 6% of Microsoft tests in one analysis and about 10% of certain triggered LinkedIn tests hit an SRM (9). Those are company-specific rates, not industry estimates, but they indicate that this is common rather than exotic.

Novelty and primacy effects are the other trap. Novelty is temporary extra engagement with something new that fades; primacy is initially poor performance that improves as people learn the change (10). Both are real. But a widely cited analysis of puzzling experiment outcomes found that most suspected novelty and primacy patterns are not real effects at all, just the high statistical variability of cumulative results in the first few days (11). Run for complete behavioural cycles, report effects by time since assignment, and do not invent a novelty narrative every time an early graph wobbles.

Data quality underneath all of it. Freshness measures when data was last updated; completeness assesses whether it contains what the purpose requires (12). Both should be monitored as dimensions with defined objectives rather than assumed.

The one that still bothers me

The most uncomfortable data problem I met in ten years was not sinister. It was arithmetic.

One of the biggest cosmetics groups in its market had never reconciled customer identity between its offline systems and its several online storefronts. When a customer bought through a channel they had not used before, the system stored them as a new person. Category tracking was broken badly enough that the “most purchased category” field they sent us contained values that meant nothing at all. They eventually paid a CRM vendor a great deal of money to clean it up.

Every personalisation number that business produced before that clean-up was counting something other than what it claimed to count. Not wrong by a few percent. Counting a different thing.

Before you invest in better models, confirm that your identity resolution, event capture and attribute coverage do what you believe they do. Adobe, for instance, surfaces the percentage of profiles containing a value for a given attribute and flags attributes populated in under 25% of profiles (13). That kind of check costs nothing and occasionally saves a year.

Reporting cadence and ownership

Match the cadence to the behaviour, not to the calendar.

Operational metrics, meaning latency, failure rates, coverage and delivery, can be monitored continuously and should alert. Experiment results should be read at predefined checkpoints, not watched daily, because watching daily is how teams stop tests the moment they look good. Commercial and relationship outcomes need fixed observation windows long enough for returns, cancellations and repeat purchases to mature.

Ownership needs to be explicit in three places: who owns the metric definition, who owns the data pipeline behind it, and who is accountable for the decision the metric supports. When those three are the same person, results get optimistic. When nobody holds the first one, the same KPI name acquires three meanings within a year.

Two published examples are worth copying. Booking.com described building a searchable central repository containing both successful and unsuccessful experiments, with transparent monitoring of data-pipeline reliability and standardised recording of hypotheses, segments and final decisions (14). Spotify described separating variant creation, assignment, pre-launch QA, overlap coordination and automated validation of proposed tests into distinct responsibilities (15).

Neither is about a clever metric. Both are about making results trustworthy and repeatable, which is the part that actually decays without maintenance.

Turning findings into decisions

A result becomes a decision when three things are on the same page.

The absolute incremental effect, with its uncertainty. Not the relative lift alone.

The contribution after costs. Incentive spend, delivery cost, service cost and returns come off before anyone celebrates. Revenue is not profit, and a personalised experience that shifts volume into discounted lines can lift revenue and reduce margin.

Any guardrail that moved the wrong way. An experience that raises short-term revenue per visitor while raising opt-outs and complaints may be moving value from the future into the present.

And statistical significance is not commercial importance. A significant result on a tiny effect in a large sample is a real finding about a difference too small to act on.

The mistakes I saw most often

Calling an exposed-versus-unexposed comparison incrementality. The most common measurement error in personalisation, and the most expensive.

Reading one statistic in isolation. Conversion without order value, order value without conversion, clicks without anything downstream.

Stopping a test when it looks good. Predefine the duration and the decision rule.

Changing a definition mid-flight. Engagement, eligibility, attribution window. Any of them invalidates the comparison.

Treating a result as permanent. A measured lift describes a population in a context at a time. Audience mix shifts, competitors change expectations, novelty fades. Attach a date and a review point to every result, and either keep a smaller holdout running or re-test periodically rather than letting an eighteen-month-old number keep earning credit.

Importing someone else’s benchmark. Definitions, populations, channels, purchase cycles, margins and exposure rules vary too much. Your own historical baseline under consistent definitions is worth more than any industry figure, including the ones in this article.

A personalisation measurement checklist

Before a test launches, confirm:

  1. the decision this measurement supports is written down;
  2. eligibility is defined and recorded before treatment;
  3. assignment is random, at customer level, and persistent across channels;
  4. exposure is logged separately from assignment;
  5. one primary metric is named, with diagnostics and guardrails around it;
  6. the observation window covers the behaviour you care about;
  7. definitions of engagement, conversion and attribution are frozen for the duration;
  8. instrumentation checks, including sample ratio, will run before anyone reads a result;
  9. the analysis plan specifies intention-to-treat;
  10. costs are included in the commercial readout;
  11. there is a stated condition for stopping or retiring the experience;
  12. someone owns the definition, someone owns the pipeline, and someone owns the decision.

Personalisation measurement is not difficult because the statistics are hard. It is difficult because the comparison is easy to get wrong and the wrong comparison produces a number that looks exactly like a right one.

Frequently Asked Questions

How do you prove personalisation is working rather than just looking like it works?

Randomly assign eligible customers to receive the personalised experience or not, hold that assignment persistently across every channel, and compare the assigned groups rather than the exposed ones. Without a control group you are measuring the behaviour of people who were already more likely to convert, since customers are unexposed for reasons that independently predict their behaviour. A study of Amazon browsing covering 2.1 million users estimated that at least 75% of recommendation click-through activity would have occurred without the recommendations. That gap is invisible in a dashboard and only appears against a holdout.

Why is comparing exposed and unexposed customers unreliable?

Because the two groups differ for reasons that independently predict the outcome. Customers are unexposed because they lacked consent, were unidentified, arrived through a different channel, or were less interested to begin with, and customers who clicked a recommendation had already demonstrated intent. Conditioning on exposure selects people partly on the thing you are trying to measure. An instrumental-variable study of Amazon browsing covering 2.1 million users estimated that at least 75% of recommendation click-through would have occurred without the recommendations. The fix is to define eligibility before treatment and analyse by assignment.

How long should a personalisation experiment run?

Long enough to cover complete behavioural cycles for your category and to let the outcome you care about actually appear, which for returns, cancellations or repeat purchases can be considerably longer than the test itself. Predefine the duration before launch. Watching results daily and stopping when they look good produces systematically overstated wins. Report effects by time since assignment so you can distinguish a sustained effect from early statistical noise, which accounts for most apparent novelty patterns.

Why do experiment results not add up to the programme's total impact?

Because you ship the winners. If you launch only experiments that came back significantly positive and then sum their estimates, the total overstates the programme, since the selection itself biases the estimates upward. Airbnb documented this as a winner's curse problem and described using a persistent split holdout as a meta-experiment to measure the combined effect directly. Seasonality, interactions between treatments and differences between short- and long-term effects also contribute.

What should you check before trusting any personalisation result?

Instrumentation, before statistics. Check that the observed split between variants matches the configured split, since a sample ratio mismatch signals problems in bucketing, identifiers, logging, joins or analysis that can reverse a decision. Confirm identity resolution actually works, because a business that stores a returning customer as a new person when they switch channel is counting something other than what its reports claim. Confirm no definition changed mid-test. Then look at the numbers.

References

  1. McKinsey & Company. The value of getting personalization right, or wrong, is multiplying. 12 November 2021. https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/the-value-of-getting-personalization-right-or-wrong-is-multiplying
  2. McKinsey & Company. 2021 Top picks: From recovery to growth, survey methodology exhibit. December 2021. https://www.mckinsey.com/~/media/mckinsey/business%20functions/marketing%20and%20sales/our%20insights/2021%20top%20picks%20from%20recovery%20to%20growth/2021-top-picks-from-recovery-to-growth.pdf
  3. Forrester Consulting, commissioned by Adobe. How to improve the ROI of personalization at scale in the era of AI. May 2025. https://business.adobe.com/content/dam/dx/us/en/resources/reports/personalization-at-scale-report/personalization-at-scale-report.pdf
  4. Sharma, A., Hofman, J. M. and Watts, D. J. Estimating the Causal Impact of Recommendation Systems from Observational Data. ACM Conference on Economics and Computation, 2015. https://arxiv.org/abs/1510.05569
  5. Airbnb. Selection bias in online experimentation: thinking through a method for the winner’s curse in A/B testing. 29 May 2017. https://medium.com/airbnb-engineering/selection-bias-in-online-experimentation-c3d67795cceb
  6. Netflix Technology Blog. Recommending for Long-Term Member Satisfaction at Netflix. 29 August 2024. https://netflixtechblog.com/recommending-for-long-term-member-satisfaction-at-netflix-ac15cada49ef
  7. Learning the Covariance of Treatment Effects Across Many Weak Experiments. ACM SIGKDD, 2024. https://arxiv.org/html/2402.17637v2
  8. Fabijan, A. et al. Diagnosing Sample Ratio Mismatch in Online Controlled Experiments: A Taxonomy and Rules of Thumb for Practitioners. ACM SIGKDD, July 2019. https://www.microsoft.com/en-us/research/publication/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonomy-and-rules-of-thumb-for-practitioners/
  9. Microsoft Research. Diagnosing Sample Ratio Mismatch in A/B Testing. 14 September 2020. https://www.microsoft.com/en-us/research/articles/diagnosing-sample-ratio-mismatch-in-a-b-testing/
  10. Microsoft. Novelty and Primacy: A Long-Term Estimator for Online Experiments. February 2021. https://arxiv.org/pdf/2102.12893
  11. Kohavi, R. et al. Trustworthy Online Controlled Experiments: Five Puzzling Outcomes Explained. ACM SIGKDD, August 2012. https://ai.stanford.edu/~ronnyk/puzzlingOutcomesInControlledExperiments.pdf
  12. Google Cloud. Auto data quality overview. Updated 24 July 2026. https://docs.cloud.google.com/dataplex/docs/auto-data-quality-overview
  13. Adobe. Segment Builder UI guide, Adobe Experience Platform documentation. Updated 18 June 2026. https://experienceleague.adobe.com/en/docs/experience-platform/segmentation/ui/segment-builder
  14. Democratizing online controlled experiments at Booking.com. 23 October 2017. https://arxiv.org/abs/1710.08217
  15. Spotify Engineering. Experimenting at Scale, the Spotify Home Way. 15 June 2023. https://engineering.atspotify.com/2023/06/experimenting-at-scale-the-spotify-home-way
◍ herm · cite this

Use this guide as a source

If it settled an argument in your reporting, cite it, and add Herm as a preferred source so the next answer you get from search or a model comes from work with its definitions attached.

└ Erul, İ. (2026) How to Measure the Effectiveness of Personalisation. Herm. www.herm.io/blog/how-to-measure-the-effectiveness-of-personalization/
İlkem Erul
Written by

İlkem Erul

Contributor

I have over nine years of experience in digital marketing, account management, and B2C loyalty. I've helped global brands grow, and now, as a co-founder of Herm.io, I work on smarter, safer shopping experiences for consumers.

More from İlkem →

Related reading

All in this category →

More in Measuring Marketing Performance

01 Personalisation KPIs: 19 Metrics, Formulas & Definitions 02 Marketing Metrics by Funnel Stage: Formulas & KPIs 03 How to Measure Marketing Performance: A Practical Framework

Get the next guide

Readiness

Attribution you can't defend is one symptom. See how five AI models currently describe, price and recommend your brand.

Get your score