Measuring Performance

How to Measure Marketing Performance

Build a marketing measurement framework covering KPI hierarchies, definitions, attribution, experiments, MMM and decision-making.

How to Measure Marketing Performance

A marketing measurement framework is not a dashboard. It is a set of agreements about which decisions get made with which evidence, and it fails at the agreement stage far more often than at the analytics stage.

This guide covers how to build one: translating objectives into measurement questions, constructing a KPI hierarchy, defining metrics so they mean the same thing to everyone, choosing between attribution, experiments and marketing-mix modelling, and running the review process that turns findings into decisions.

It does not define individual metrics. For formulas, denominators, required inputs and interpretation warnings across the funnel, see our reference on marketing metrics and KPIs by funnel stage.

Start with the decisions, not the metrics

The IAB’s guidance on modernising marketing-mix modelling opens with the same instruction, and it applies well beyond MMM: identify the decisions, the possible actions and the desired outcomes first, then agree the KPIs, lookback windows, thresholds and escalation rules (1).

Most measurement frameworks are built the other way round. Someone lists the metrics available, arranges them by channel, and calls the result a framework. It produces reports nobody acts on, because no decision was ever attached to any number.

There are only a handful of decisions marketing measurement supports:

  • how much to spend in total;
  • how to divide that spend;
  • whether to continue, change or stop a specific activity;
  • whether a change worked;
  • whether something is broken.

Write down which of those each report serves. Any metric that serves none of them is a diagnostic at best, and probably belongs somewhere other than the front page.

Turn objectives into measurement questions

An objective is a direction. A measurement question is answerable.

“Grow the business” becomes “is our acquisition of new customers increasing at a cost that our contribution margin supports?” That question names a population, a metric family, a comparison and an economic constraint, which means someone can go and answer it.

Do this before choosing metrics, because the question determines the denominator. “Are we acquiring more customers?” and “are we acquiring more efficiently?” require different calculations from the same data, and teams that skip this step end up arguing about which one the number was.

Build a hierarchy, not a list

Not every metric is a KPI. A hierarchy has three levels and the relationship between them should be explicit.

Business outcomes sit at the top: revenue, contribution margin, customer acquisition, retention. These are what the company is actually trying to move.

Marketing outcomes sit beneath: acquisition volume, cost per new customer, retained revenue, category demand. These are what marketing plausibly influences.

Channel diagnostics sit at the bottom: click-through rate, cost per click, engagement rate, deliverability. These explain movement in the level above. They are not objectives.

The common failure is promoting a diagnostic. A team that optimises click-through rate as a goal will get a higher click-through rate, and there is no guarantee that anything above it moves.

Financial and customer outcomes belong above channel diagnostics, always. If your reporting inverts that, the reporting is telling people what to care about, and it is telling them wrong.

Define metrics so they survive contact with other teams

This is the least glamorous part of the framework and it prevents more damage than anything else in this article.

For every metric that reaches a report, record the name, the exact formula, the numerator, the denominator, the included population, the exclusions, the data source, the attribution rule, the reporting window, the owner and the known limitations.

The MRC’s outcome standards require that measures be relevant and aligned to campaign goals, supported by consistent definitions, and that calculated efficiency measures disclose their sources and calculations (2). That is the standard worth holding internally, whether or not you are subject to it.

Two specific disciplines matter more than the rest.

Denominators have to be stated. Every metric needs an explicit event, population and denominator. Shared labels such as conversion rate frequently conceal incompatible definitions.

Cross-platform totals cannot simply be added. The MRC’s cross-media standards exist because comparable presentation requires harmonised viewability bases, duration, invalid-traffic filtration and audience deduplication (3). Summing reach across platforms double-counts everyone exposed in more than one place.

What the reported number was, and what it described

A French luxury house I worked with had a clear view of its customer: women aged 40 and over. That was the ideal customer profile, it appeared in every deck, and the digital targeting was built on it.

When we actually looked at who bought on the website, the answer was women between 27 and 40 and men between 25 and 35. The 40-plus customer was real. She was the boutique clientele, and the ICP had been built from the offline heritage channel and then applied to a digital business that had a different audience.

Nobody had lied. The definition had simply never been checked against the population it was being used to describe. Every performance number computed against that target was accurate and describing the wrong people.

This is why the definition step is not administrative. A metric inherits whatever assumption is buried in its population.

Identify data sources and ownership

Three roles need naming for every KPI, and they should rarely be the same person.

The definition owner decides what the metric means and approves changes to it.

The pipeline owner is accountable for the data behind it being correct and available.

The decision owner is accountable for acting on it.

The WFA’s cross-media framework asks for neutral governance representing both buying and selling parties, transparent definitions, disclosure of bias and error, and audits of data and methods (4). Internally the equivalent is simply making sure the person who reports a number is not the only person who can change how it is calculated.

Establish baselines and targets

A baseline is the same metric, computed the same way, on the same population, over a period long enough to contain your normal variation. Anything else is a comparison against a moment.

Targets need a stated basis. Four defensible starting points are your own historical performance under consistent definitions, a contemporaneous control, expected economics derived from margin and customer lifecycle, and a documented service objective. An external benchmark can occasionally serve, but only where the definitions, market, channel and sample genuinely match the decision, which is uncommon.

Benchmarks travel badly, and the research is unusually clear about how badly. A widely repeated figure putting luxury fashion ecommerce conversion at 0.5 to 0.8% traces back through vendor content to a source that cannot be inspected. A transparently methodologised benchmark covering more than seven billion sessions across 400 sites put luxury conversion at 1.01% (5). The point is not that one number is right. It is that the same category produced materially different figures under different samples, periods, device mixes and conversion definitions, and neither tells you what your business should expect.

The peaks tell you who your customers are

There is one diagnostic I run on every new account before looking at anything else, and it takes about five minutes. Open the analytics, chart conversions over a full year, and look at where the peaks are.

If the peaks appear only during sale periods, with no smaller lifts around new-season launches or product releases, the customer base is discount-driven. That is not a campaign finding, it is a structural one, and it changes what every subsequent number means.

It also explains why reflexive discount-matching is so destructive. A brand that mirrors a competitor’s promotion every time teaches its own customers to abandon their normal buying cycle and wait for the next reaction. Sometimes you match the offer and lose the customer anyway, because by then the only thing they are shopping on is price.

Choose between attribution, experiments and marketing-mix modelling

These three methods answer different questions. Treating them as competing versions of the same answer is the most consequential methodological error in marketing measurement.

Attribution allocates credit; it does not establish cause

Last-click assigns all credit to the final interaction. Data-driven models redistribute credit by comparing converting and non-converting paths (6). Both are allocation rules operating on observed data. Neither estimates what would have happened otherwise.

Google’s own research identifies the constraint plainly: attribution accuracy is limited by the model’s assumptions and by the quality and completeness of available data, and conventional attribution can miss conversions created through upstream effects like additional visits, brand searches or awareness (7).

The strongest evidence of how far apart attributed and incremental can sit comes from eBay. When brand-keyword advertising was suspended, almost all the forgone paid clicks and attributed sales simply moved to organic search. For non-brand search, conventional regressions produced ROI estimates above 1,400% with controls, while the experimental estimate was −63% (8).

Those numbers belong to eBay’s circumstances and should not be generalised. The lesson that generalises is that platform-attributed purchases and incremental purchases are different quantities, and the difference can be larger than the entire measured effect.

A peer-reviewed comparison against 15 randomised Facebook experiments covering 500 million user-experiment observations found the same thing more systematically: even after conditioning on extensive demographic and behavioural variables, observational methods frequently failed to recover the randomised results (9).

Use attribution for: operational reporting, journey diagnostics, in-platform optimisation. Do not use it for: claiming incremental effect.

Experiments estimate the counterfactual, and have their own limits

Randomised user or geo experiments are the right tool for causal questions. Geo-experiments assign geographic units to treatment and control where user-level randomisation is unavailable, though advertisers often have few and heterogeneous geographic units, and budget constraints can cause interference between them (10).

Be honest about statistical reality too. Across 25 large field experiments representing $2.8 million in digital advertising spend, the median ROI confidence interval exceeded 100 percentage points, because advertising effects are frequently small relative to the volatility of individual purchasing (11). Properly powered tests can require samples large enough to make some questions expensive or infeasible.

A wide confidence interval is a finding. It is not the same as no effect, and it should not be reported as one.

MMM answers aggregate allocation questions, conditional on assumptions

Google’s Meridian documentation describes MMM as fundamentally a causal problem, addressed through Bayesian modelling and explicit causal assumptions (12). That framing is correct and it carries an obligation: the results are causal only to the extent the assumptions, controls and specification are credible.

Google is equally explicit that model fit does not validate causal accuracy, and that directly validating causal inference requires well-designed experiments (13). Meta’s Robyn documentation makes the complementary recommendation, that experimental results be used to calibrate MMM, while warning that an experiment measures marginal impact for a particular period rather than the whole modelling window (14).

Use MMM for: cross-channel contribution, longer horizons, online and offline together, response curves, budget scenarios. Calibrate it against experiments. Do not treat good fit as proof.

Triangulate carefully or not at all

Comparing methods is only meaningful after aligning the outcome, time window, geography, treatment definition and estimand. A channel experiment from one quarter does not validate a multi-year MMM coefficient, and presenting it as though it does manufactures false confidence.

Account for short-term and long-term effects

Binet and Field’s IPA Databank analysis is the standard reference here, and it is usually cited badly.

The work covers 996 campaigns, 700 brands and 83 categories, and reports an optimum activation share of roughly 40% across the analysed cases, which is the origin of the widely quoted 60:40 split (15). Two qualifications matter. The Databank collates information from questionnaires submitted as part of IPA Effectiveness Award entries, meaning it describes submitted effectiveness cases rather than a random sample of all campaigns (16). And the IPA’s own later work states that the optimal mix differs by brand type and context (17).

So: effective campaigns generally need both short-term activation and longer-term brand building, and the balance is a starting hypothesis rather than a budgeting rule. Any framework that measures only what converted this month will systematically underinvest in the half of the effect it cannot see.

Design reports people act on

Three rules.

Report the numerator, the denominator, the rate, the absolute difference, the relative difference and the window. A percentage without its underlying volume misleads whoever reads it, including the person who calculated it.

Separate significance from importance. A statistically significant result on a tiny effect in a large sample is a real finding about a difference too small to act on.

Say what you do not know. In the February 2019 CMO Survey, 36.4% of respondents said they proved the long-term impact of marketing spend quantitatively, 50.7% reported a qualitative sense without quantitative impact, and 12.9% said they had not been able to show it yet (18). That is a normal distribution of certainty. A report that presents everything at the same confidence level is hiding something.

Diagnose conflicting metrics

When two numbers disagree, work through this order: are the definitions the same, are the populations the same, are the windows the same, are the attribution rules the same, and only then, is one of them wrong.

Most conflicts resolve in the first three steps.

The ceiling on what marketing measurement can tell you

There is a limit to this that vendors do not advertise and that I spent years running into.

I watched clients underperform for reasons that no campaign could offset. Late deliveries. Wrong data feeding stock-management systems, so customers were shown products that were not really available. Storage and logistics failures that produced dissatisfaction long before any marketing message reached anyone. We had the data to demonstrate exactly why those brands were not performing, and no ability whatsoever to influence brand image, pricing strategy or operations.

That is worth stating plainly in a measurement framework, because it prevents a specific mistake. When a metric is stubborn, the framework’s job is to tell you where the problem is, not to assume the problem is inside marketing. A conversion rate that will not move may be reporting a fulfilment problem. Marketing measurement can locate the constraint. It cannot always fix it, and a framework that only looks at marketing inputs will keep proposing marketing solutions to problems that are not marketing problems.

Account for emerging measurement blind spots

Part of product discovery and, increasingly, transactions now happen where conventional analytics cannot observe them.

Structured merchant feeds and in-assistant commerce are operational rather than speculative. OpenAI publishes a product-feed specification (19), Google and Shopify have co-developed an open commerce protocol (20), and Microsoft began rolling out an in-conversation checkout in January 2026 that completes a purchase without a redirect (21).

The measurement consequence is structural. Google Analytics has no durable default channel for AI assistants; you build one from source rules and maintain it (22). Search Console folds AI Overviews and AI Mode clicks into overall Web search with no separate breakdown (23). And where a purchase completes inside an assistant, there may be no browser session at your site at all.

Be sceptical of the performance claims circulating alongside this. Adobe reports AI-referred retail visitors converting substantially better than all non-AI sources combined, from very large telemetry (24), but the comparison group is all other traffic rather than organic search, visitors arriving after extensive assistant research are heavily selected, and Adobe sells AI visibility products. It is a selection-adjusted funnel observation, not evidence of incremental demand.

Treat assistant exposure, referred sessions and transactions as three separate layers. Conventional analytics reliably sees only the middle one. Build that gap into the framework as a known unknown rather than waiting for a figure that does not yet exist.

Improve the framework over time

A framework decays. Definitions drift, pipelines change, and results expire.

Schedule three reviews. Definitions annually, or whenever a data source changes. Pipeline integrity continuously, with alerts. And results themselves on a stated cadence, because a measured effect describes a population in a context at a time. Audience mix shifts, competitors change what customers expect, and a number calculated eighteen months ago should not still be earning credit unchallenged.

For personalisation programmes specifically, where eligibility and exposure create their own measurement problems, see our guide to measuring personalisation effectiveness.

Frequently Asked Questions

What is the difference between a metric and a KPI?

A KPI is a metric that a decision depends on. Everything else is a diagnostic. A useful hierarchy has three levels: business outcomes such as revenue and contribution margin at the top, marketing outcomes such as acquisition cost and retained revenue beneath them, and channel diagnostics such as click-through rate and cost per click at the bottom. Diagnostics explain movement in the level above. The common failure is promoting one to a goal, because a team that optimises click-through rate will get a higher click-through rate with no guarantee that anything above it moves.

Should we use attribution, experiments or marketing-mix modelling?

They answer different questions. Attribution allocates credit among observed touchpoints and is suited to operational reporting and in-platform optimisation, but it does not estimate what would have happened otherwise. Experiments estimate that counterfactual and are the right tool for causal questions, though effects are often small relative to purchasing volatility and properly powered tests can be expensive. Marketing-mix modelling handles aggregate cross-channel allocation over longer horizons, and both Google and Meta recommend calibrating it against experimental results rather than trusting model fit.

Why do platform-reported conversions overstate marketing's impact?

Because assigned credit and incremental effect are different quantities. When eBay suspended brand-keyword advertising, almost all forgone paid clicks and attributed sales moved to organic search, and for non-brand search a conventional regression produced an ROI estimate above 1,400% while the experimental estimate was negative. A peer-reviewed comparison against 15 randomised experiments at Facebook found observational methods frequently failed to recover randomised results even after extensive controls. Platforms also claim the same sale independently, so summing across them double-counts.

How should we set marketing targets without industry benchmarks?

There are four defensible bases: your own historical performance under consistent definitions, a contemporaneous control group, expected economics derived from your margin and customer lifecycle, or a documented service objective. Industry averages travel badly. A widely repeated luxury ecommerce conversion figure of 0.5 to 0.8% traces to a source that cannot be inspected, while a transparently methodologised benchmark across seven billion sessions put the same category at 1.01%. Different samples, periods and conversion definitions produce materially different numbers for the same category.

How should a measurement framework handle channels it cannot see?

Name the gap rather than ignoring it or filling it with a vendor estimate. Some discovery and, increasingly, some transactions now occur inside AI assistants, where conventional analytics has no session to record. Google Analytics has no durable default channel for assistants and Search Console folds AI features into overall Web search. The framework response is to separate exposure, referred sessions and transactions into distinct layers, document which layers your instrumentation can actually observe, and treat the unobserved portion as a known unknown in reporting rather than assuming it is zero.

References

  1. Interactive Advertising Bureau. Modernizing MMM: Best Practices for Marketers. December 2025. https://www.iab.com/wp-content/uploads/2025/12/IAB_Modernizing_MMM_Best_Practices_for_Marketers_December_2025.pdf
  2. Media Rating Council. MRC Outcomes and Data Quality Standards. September 2022. https://mediaratingcouncil.org/sites/default/files/Standards/MRC%20Outcomes%20and%20Data%20Quality%20Standards%20%28Final%29.pdf
  3. Media Rating Council. MRC Cross-Media Audience Measurement Standards (Phase I Video). September 2019. https://mediaratingcouncil.org/sites/default/files/Standards/MRC%20Cross-Media%20Audience%20Measurement%20Standards%20%28Phase%20I%20Video%29%20Final.pdf
  4. World Federation of Advertisers. Establishing Principles for a New Approach to Cross-Media Measurement: An Industry Framework. https://wfanet.org/knowledge/item/2020/09/17/Global-advertisers-unveil-a-collaborative-new-approach-to-cross-media-measurement
  5. Contentsquare. 2020 Digital Experience Benchmark. 2020. https://go.contentsquare.com/hubfs/eBooks/2020%20Digital%20Experience%20Benchmark/2020%20Digital%20Experience%20Benchmark%20Report%20PDF%20English.pdf
  6. Google Ads Help. About data-driven attribution. Accessed 27 July 2026. https://support.google.com/google-ads/answer/6394265
  7. Sapp, S. and Vaver, J. Toward Improving Digital Attribution Model Accuracy. Google Research, 2016. https://research.google/pubs/toward-improving-digital-attribution-model-accuracy/
  8. Blake, T., Nosko, C. and Tadelis, S. Consumer Heterogeneity and Paid Search Effectiveness: A Large Scale Field Experiment. 2014. https://faculty.haas.berkeley.edu/stadelis/BNT_ECMA_rev.pdf
  9. Gordon, B. R., Zettelmeyer, F., Bhargava, N. and Chapsky, D. A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook. Marketing Science, 2019. https://pubsonline.informs.org/doi/10.1287/mksc.2018.1135
  10. Chen, A. and Au, T. Robust Causal Inference for Incremental Return on Ad Spend with Randomized Paired Geo Experiments. Annals of Applied Statistics, 2022. https://research.google/pubs/robust-causal-inference-for-incremental-return-on-ad-spend-with-randomized-paired-geo-experiments/
  11. Lewis, R. A. and Rao, J. M. The Unfavorable Economics of Measuring the Returns to Advertising. Quarterly Journal of Economics, November 2015. https://academic.oup.com/qje/article-abstract/130/4/1941/1914592
  12. Google. Introduction to Bayesian Modeling and Causal Inference Theory, Meridian documentation. Updated 15 May 2026. https://developers.google.com/meridian/docs/causal-inference/intro
  13. Google. Assess the model fit and the results, Meridian documentation. Updated 30 June 2026. https://developers.google.com/meridian/docs/post-modeling/model-fit
  14. Meta Platforms. An Analyst’s Guide to MMM, Robyn documentation. Accessed 27 July 2026. https://facebookexperimental.github.io/Robyn/docs/analysts-guide-to-MMM/
  15. Institute of Practitioners in Advertising. The Long and the Short of It: 10 Key Principles of Success. 11 September 2013. https://ipa.co.uk/media/14463/long_and_short_of_it_presentation_final_030424.pdf
  16. Grande, C. Six things we learnt about “The Long and the Short of It”, 10 years on. IPA, 18 October 2023. https://ipa.co.uk/knowledge/ipa-blog/six-things-we-learnt-about-the-long-and-the-short-of-it-10-years-on
  17. Institute of Practitioners in Advertising. Les Binet and Peter Field. Updated 12 November 2025. https://ipa.co.uk/knowledge/effectiveness-research-analysis/les-binet-peter-field
  18. The CMO Survey. Topline Report, February 2019. https://cmosurvey.org/wp-content/uploads/2024/03/The_CMO_Survey-Topline_Report-Feb-2019-20240328-142526.pdf
  19. OpenAI. Powering Product Discovery in ChatGPT. 24 March 2026. https://openai.com/index/powering-product-discovery-in-chatgpt/
  20. Google Developers. Under the Hood: Universal Commerce Protocol. 11 January 2026. https://developers.googleblog.com/under-the-hood-universal-commerce-protocol-ucp/
  21. Microsoft Advertising. Conversations that Convert: Copilot Checkout and Brand Agents. 8 January 2026. https://about.ads.microsoft.com/en/blog/post/january-2026/conversations-that-convert-copilot-checkout-and-brand-agents
  22. Google Analytics Help. Custom channel groups. Accessed 27 July 2026. https://support.google.com/analytics/answer/13051316
  23. Google Search Central. AI features and your website. Updated 10 December 2025. https://developers.google.com/search/docs/appearance/ai-features
  24. Adobe. AI traffic to travel sites soars nearly 200% as AI visitors now outperform traditional sources, Adobe Digital Insights. 17 June 2026. https://business.adobe.com/blog/adobe-report-ai-traffic-travel-sites-surges-200-percent
◍ herm · cite this

Use this guide as a source

If it settled an argument in your reporting, cite it, and add Herm as a preferred source so the next answer you get from search or a model comes from work with its definitions attached.

└ Erul, İ. (2026) How to Measure Marketing Performance. Herm. www.herm.io/blog/measuring-marketing-performance-unlocking-data-driven-success/
İlkem Erul
Written by

İlkem Erul

Contributor

I have over nine years of experience in digital marketing, account management, and B2C loyalty. I've helped global brands grow, and now, as a co-founder of Herm.io, I work on smarter, safer shopping experiences for consumers.

More from İlkem →

Related reading

All in this category →

More in Measuring Marketing Performance

01 Personalisation KPIs: 19 Metrics, Formulas & Definitions 02 Marketing Metrics by Funnel Stage: Formulas & KPIs 03 How to Measure Personalisation Effectiveness and ROI

Get the next guide

Readiness

Attribution you can't defend is one symptom. See how five AI models currently describe, price and recommend your brand.

Get your score