A customer data platform is not a compulsory layer in every marketing stack. It is a response to a defined data and activation problem, and if you cannot name that problem in operational terms, you are not ready to buy one.
The category is usually described as software that builds and maintains a persistent, unified customer record which other systems can then use. That description is fair, but the word unified does a lot of unearned work. A platform assembles records according to rules somebody configured. It cannot guarantee that every source is accurate, that every identifier belongs to one person, or that every merge reflects the actual relationship. Vendor documentation is clearer about this than most marketing material: identity resolution in a leading platform is a set of settings and rules that the customer has to define, not an automatic property of the product (1).
So the useful question is not which platform to buy. It is this:
Which decision or activation problem cannot be solved reliably with our current CRM, warehouse, analytics and automation systems?
A CRM supports the operational management of identified relationships: accounts, contacts, opportunities, cases and communications. A CDP collects and transforms data from many sources to produce profiles, audiences, analysis and activation. The products overlap. The organisational jobs do not. For the operational side of that boundary, see the CRM operating model.
What a CDP Should Actually Produce
A well run programme creates governed data products, not an impressive record count.
Depending on the use case, the useful outputs are a profile assembled from named sources, an account or household view, computed traits, an audience with explicit eligibility and exclusion rules, a suppression state, a destination-ready payload, an activation and exposure log, and a quality report showing latency, missing fields, rejected records and uncertain matches.
Every one of those is an intermediate capability. Their value depends entirely on whether they change a decision.
A large audience is not evidence of value. A profile is not evidence of accuracy. A campaign sent is not evidence of incremental impact. Hold on to those three sentences, because most CDP business cases quietly violate all of them.
- Name the blocked decision Describe in operational terms what someone cannot currently do. If the objective is a 360-degree view, you have not specified anything anyone will act on differently.
- Test the cheaper alternatives honestly Improving the CRM, the warehouse, instrumentation, reverse ETL or a single integration solves a surprising number of these problems at a fraction of the cost.
- Specify the output, not the platform Write the required audience or profile as a specification with entity level, refresh rate, eligibility and suppression rules. That exposes what you actually need.
- Design data contracts before architecture Agree schema, meaning, ownership, quality and change management with each source owner. A connector moves data; it does not make data safe to use.
- Measure outcomes, not activity Log eligibility, delivery and exposure, then compare against something. More profiles, audiences and campaigns can coexist with no improvement at all.
Do We Need One? Thirteen Questions Before Procurement
Work through these before shortlisting a product. Most organisations discover the answer somewhere in the first five.
1. Which Decision or Activation Problem Exists?
Describe the blocked decision operationally. A service team cannot see recent product usage before responding. A suppression request reaches email but never reaches the paid-media destinations. Analysts can define an eligible audience but activation takes two weeks. Web, app and transaction events use conflicting definitions. Marketing cannot distinguish account-level from person-level eligibility.
Avoid objectives like creating a single customer view. They describe an artefact, not a change in behaviour.
Something worth noticing here is that the size of the organisation does not settle the question. I worked with one of the biggest consumer electronics companies in the world, and the blocker was not scale or budget. Before it would commit to a targeted discount on a specific product group, it needed information it simply did not hold, so we collected it directly from end users through a custom survey and used the declared answers to build the offer. The conversion contribution from the people who responded was strong enough to give the brand confidence to discount. Even at that size, the constraint was a missing input to one decision, not an abstract desire for unified data.
2. Which Source Data Are Required?
List only the sources needed for the first use case: CRM records, account and opportunity data, web or app events, orders and returns, product usage, service interactions, preference records, loyalty data, reference data.
The first-party data guide should own collection and instrumentation standards. A CDP should consume data that already has an owner, a meaning and a quality expectation.
Different sources support different decisions. Completed transactions provide evidence that a commercial event occurred, but they may be sparse, delayed or incomplete. Behavioural data can provide earlier context, but it is often noisy and does not directly reveal motivation. Rank each source according to its relevance, quality, timeliness and incremental usefulness for the defined decision, rather than assuming that one data category is inherently superior.
3. Why Can Existing Systems Not Solve It?
Test the alternatives properly. The problem may yield to a better CRM configuration, a better warehouse model, reverse ETL, improved automation, governed analytical tables, better instrumentation, or one narrow integration.
A CDP is justified only when the required combination of ingestion, profile assembly, audience computation, governance and activation cannot be delivered acceptably through what you already own.
4. Which Profile or Audience Output Is Required?
Write the output as a specification. For example: an account-level renewal-risk audience, refreshed daily, containing active customers whose contract ends within ninety days, excluding accounts with an open complaint and all contacts suppressed for the selected channel.
That is more useful than single customer view because it exposes entity level, timing, eligibility and permission requirements all at once.
5. What Latency Is Required?
Real time is not a default requirement. Choose latency from the decision: seconds for an in-session treatment, minutes for a service or fraud intervention, hours for triggered communications, daily for most lifecycle audiences, weekly for planning.
Lower latency raises architectural and operating complexity sharply. Do not pay for faster processing than the use case can exploit.
6. Which Destinations Are Required?
Name the systems that will receive an output, and check what each one can accept: API or batch support, payload limits, identifier requirements, deletion and suppression behaviour, error reporting, retry logic, audience expiry and delivery receipts.
An audience that computes perfectly and never arrives correctly is an implementation failure, and it is one that dashboards inside the platform will not show you.
7. What Identity Capability Is Required?
Define the minimum identity problem rather than buying the broadest promise. You may need login-based person matching, contact-to-account relationships, household grouping, device-to-person association, anonymous-to-known transition, source-specific identifiers, or no cross-source identity at all.
Identity rules create false merges as well as false splits. Record the evidence used, the confidence level, the precedence rules and the consequence of a wrong match. Complete matching methodology belongs in the cross-device identity resolution guide.
8. What Permission Requirements Apply?
Map purpose, channel, jurisdiction, source and restriction before any activation is built.
Data protection has to be considered at design time and throughout the lifecycle rather than bolted on afterwards (2). Purpose limitation deserves particular attention in a CDP programme, because the whole point of the architecture is to make data from one context available in another. The rules governing when further processing is compatible with the original purpose were amended by section 71 of the Data (Use and Access) Act 2025 (3), which commenced on 5 February 2026 (4). A purpose map written before that date should be reviewed rather than assumed current.
At minimum, define which purposes each field may support, whether a lawful basis has been documented, which preferences are channel-specific, how objections and withdrawals are represented, which destinations receive suppression changes, what happens when systems disagree, and how deletion and retention propagate.
Where a person objects to direct marketing, the organisation must stop sending it. The regulator’s position on suppression is more precise than it is usually quoted: the law does not expressly require a suppression list, but the regulator says organisations should use one, retaining only the minimum information needed to recognise the person and honour their preference (5). Either way, suppression is an operational control rather than an optional segment.
9. Do We Have Data and Engineering Capacity?
A composable design preserves existing data models but needs people who can operate ingestion, transformation, orchestration, testing, access control, monitoring and incident response.
A packaged platform reduces assembly work but still requires source owners, schema decisions, identity rules, quality monitoring, destination testing, permission design and a named operating owner. No architecture removes ownership work. It only moves it.
10. Who Owns Ongoing Operation?
Name accountable roles before procurement: a product owner for use cases and priorities, a data owner for meaning and permitted use, an engineering owner for pipelines and reliability, a privacy or legal owner for restrictions, a marketing operations owner for audiences and destinations, an analytics owner for measurement, and a business owner for outcome realisation.
A steering committee is not a substitute for named operational ownership.
11. What Is the Total Cost?
Include licences, usage charges, storage and compute, implementation, connectors, professional services, internal engineering, data modelling, testing, privacy and security review, training, support, monitoring, destination maintenance, migration, duplicated infrastructure and exit.
Then compare that total against the cost of the narrowest viable alternative. For tool overlap and fully loaded cost across the wider portfolio, use the martech stack audit guide.
12. What Is the Exit and Portability Plan?
Before purchase, define how profiles, event history, schemas and consent states can be exported, whether identity rules are portable, whether computed audiences can be reproduced elsewhere, which transformations are proprietary, what happens to data after termination, how connectors are replaced, the notice periods, transition support, and the cost of running old and new systems in parallel.
Portability is an architectural requirement, not a procurement footnote.
13. What Evidence Shows the Capability Will Change an Outcome?
Write the causal chain out. More reliable suppression state, therefore fewer unwanted messages, therefore fewer complaints and lower compliance risk. Or: faster product usage data, therefore earlier service intervention, therefore a higher resolution rate.
If the chain terminates at more profiles, more audiences or more campaigns, the business case is incomplete.
Choosing an Architecture
| Architecture | Strength | Dependency | Best suited to | Main risk |
|---|---|---|---|---|
| Packaged | Integrated ingestion, profile assembly and activation | The vendor platform | Teams needing a managed environment quickly | Lock-in and duplicated storage |
| Composable or warehouse-native | Uses the data platform you already own | Strong data engineering capability | Organisations with a mature warehouse | Operating complexity carried in-house |
| Hybrid | Managed speed where needed, owned models elsewhere | Multiple systems at once | Genuinely mixed requirements | Duplicated data, duplicated logic, disputed ownership |
| No CDP | Uses CRM, warehouse and narrow integrations | Existing systems | Limited or well-defined use cases | Capability gaps if requirements grow |
A composable design places owned data infrastructure at the centre and adds selected capabilities around it, which can reduce data copying and preserve existing models (6). It does not reduce the need for dependable pipelines, governance and support, and that guidance comes from a vendor with a commercial interest in warehouse-centred architecture, so read it as a description of the pattern rather than proof of its superiority.
A packaged platform can cut integration work through managed connectors, but it may duplicate data and make profile logic or activation behaviour harder to reproduce elsewhere.
A hybrid approach is rational when some capabilities need managed speed and others need warehouse control. It also carries the highest risk of two teams maintaining two versions of the same logic.
No CDP is a valid architectural decision, and it should be on the table until the evidence removes it.
Design the Data Products Before the Platform
Define Source Contracts
A data contract is an agreement between a producer and its consumers covering structure, semantics, quality, ownership and service expectations. The Open Data Contract Standard provides a vendor-neutral, machine-readable format for documenting those elements (7).
For each source, specify the producer and owner, the business meaning, schema and data types, allowed values, keys and identifiers, event time and processing time, the freshness target, completeness and validity rules, retention, permitted uses, the breaking-change process, a support contact and the failure response.
A connector is not a contract. It moves data. It does not make the data understandable or safe to use.
Define the Entity Model
Do not force every use case into one customer object. Define person, contact, account, household, device, subscription and product separately, then define the relationships between them.
A person may have several contacts. A contact may relate to several accounts. A household is not a person. A device may be shared. An account may contain several decision-makers who disagree. Each of those distinctions changes permission, routing, analysis and activation.
Treating data as a product with a clear schema, defined lifecycle and semantic consistency is a recognised architectural principle rather than an invention of this article (8), though it comes from platform vendors and describes a pattern rather than an industry standard.
Define Profile and Audience Logic
Every computed profile or audience should carry an owner, a purpose, an entity level, source dependencies, inclusion rules, exclusion rules, permission logic, freshness, expiry, known limitations, test cases and a version.
Treat important audiences as governed products, not as disposable query results that somebody rebuilds slightly differently next quarter.
An Implementation Sequence
Define use cases and explicit non-goals. Inventory the specific tables, events, APIs and files the pilot needs. Agree data contracts with producers. Document person, account, household and device concepts, including prohibited merges. Build a source-to-purpose-to-destination permission map. Select the architecture against use case, capability, operating capacity, total cost and portability. Design ingestion, specifying batch or streaming method, retries, ordering, late events, replay, deletion, encryption and observability. Build testable transformations with lineage from source to output. Version the identity, trait, eligibility, suppression and expiry rules.
Then establish quality monitoring across contract failures, freshness, completeness, invalid values, duplicated profiles, uncertain matches and unexpected audience movement. Connect destinations and test payload mapping, delivery, retries, rejection, expiry and deletion in each one. Test identity and suppression with positive, negative and edge-case records, including shared email addresses, recycled phone numbers, account changes, anonymous-to-known transitions, withdrawals and conflicting preferences.
Only then pilot one use case with named operators. Log activation and exposure: who was eligible, what was sent, where it was delivered, when exposure occurred and which version of the logic applied. Measure incremental outcomes using random assignment or another defensible comparison, because a platform report showing conversions among exposed customers is not incremental evidence. Calculate the real operating cost. Then scale, revise or stop.
For transaction inputs and reconciliation across commerce environments, use the cross-platform purchase data guide. For downstream model and treatment selection, use the AI personalisation guide.
What to Measure
Track the platform as an operating system and the use case as a decision system. They fail independently.
| Area | Measures |
|---|---|
| Source coverage | Required sources connected and producing contracted data |
| Data latency | Source-to-usable-output time by use case |
| Contract reliability | Failed schemas, missing fields, invalid values, breaking changes |
| Profile quality | Duplicate rate, merge and split errors, unresolved records |
| Identity confidence | Match evidence, confidence bands, manual-review outcomes |
| Audience eligibility | Correct inclusion, exclusion, expiry and entity level |
| Suppression accuracy | Conflicts, propagation delay, destination failures |
| Destination delivery | Accepted, rejected, retried and expired records |
| Exposure logging | Eligibility, delivery and treatment evidence |
| Cost | Implementation, licence, compute, storage, operating labour |
| Outcome | Incremental commercial, service, risk or customer result |
Calculating Value Without Pretending There Is a Universal Return
You will be shown a return figure during procurement. Understand what it is before you repeat it internally.
A vendor-commissioned economic impact study published in August 2024 modelled a 186 per cent three-year return and payback inside six months for one major CDP (9). That study was commissioned by the vendor, built from a single interviewed organisation, applied attribution and financial assumptions and risk adjustments, and tells readers to substitute their own estimates. It is a worked example of a method. It is not a category benchmark, and the older figure that gave this page its original headline was the same kind of artefact.
Build your own model instead.
On the cost side, total cost equals platform plus implementation plus data plus engineering plus operations plus governance plus destinations plus exit.
On the benefit side, count only outcomes you can evidence: manual work that is actually removed or redeployed, systems that are actually decommissioned, reduced complaints or permission errors, measurably faster service or sales decisions, incremental profit from a tested treatment, reduced destination waste demonstrated against a comparison, or avoided risk with a documented basis.
Do not count the same benefit twice. Do not treat time saved as cash unless capacity or cost genuinely changes. Do not attribute all downstream uplift to the data layer.
Common Failure Modes
Buying before defining the decision turns the project into an open-ended integration programme with no completion criteria. Treating every source as equally valuable increases cost, risk and model complexity for data nobody uses. Building identity rules without consequence testing produces merges that expose data, misroute communications and distort reporting. Confusing a profile with truth forgets that a profile is a versioned interpretation assembled from sources and rules. Failing to propagate suppression makes a centrally stored preference worthless. Measuring activity instead of outcomes lets record counts grow while nothing improves. Underfunding operation ignores that pipelines, mappings, destination APIs, permissions and definitions all change after launch.
There is one more, and it is the one I would put above the others because no amount of data engineering touches it. A data platform cannot fix a brand problem, a pricing problem or an operations problem. I watched clients with excellent data underperform because deliveries arrived late, because stock systems were fed wrong information, or because their pricing was simply uncompetitive. Even when we could collect everything needed to prove exactly why they were struggling, nobody in the marketing technology layer had any say over brand image or pricing. This is a commercial observation rather than a technical one, but it belongs in a CDP business case: if the blocked decision sits outside marketing’s control, better data will describe the problem more precisely without moving it.
The Final Decision
Implement a CDP only when a defined decision requires data from multiple sources, current systems cannot meet the profile, audience, latency, governance or activation requirement, the organisation can operate the chosen architecture, permissions can be propagated reliably, total cost and exit are understood, and the resulting capability is likely to change a measurable customer or business outcome.
Otherwise, improve the CRM, the warehouse, the instrumentation, the automation or one narrow integration first. Those projects are cheaper, faster and much easier to reverse.
Frequently Asked Questions
What is the difference between a CDP and a CRM?
A CRM supports the operational management of identified relationships: accounts, contacts, opportunities, cases and communications, with users acting on records. A CDP collects and transforms data from many sources to produce profiles, audiences, analysis and activation. The products overlap in what they can technically store, but the organisational jobs are different, and buying one to do the other's work is a common and expensive mistake.
Does every company need a customer data platform?
No. A CDP answers a specific problem: a decision or activation that needs data from multiple sources and cannot be delivered acceptably through the existing CRM, warehouse, instrumentation or automation. If you cannot describe the blocked decision in operational terms, the honest answer is that you are not ready to buy one. No CDP is a legitimate architectural choice.
What is a composable CDP?
An architecture that keeps the data in your existing warehouse or lakehouse and adds selected capabilities such as identity resolution, audience computation and activation around it, rather than copying everything into a separate vendor platform. It can reduce data duplication and preserve your existing models, but it moves the operating burden in-house and needs engineers who can run ingestion, transformation, testing, monitoring and incident response.
Does a CDP give you an accurate single customer view?
It gives you an assembled view based on rules somebody configured. Identity resolution in commercial platforms is a set of settings the customer defines, not an automatic property of the software. Those rules can produce false merges as well as false splits, so the profile is best understood as a versioned interpretation with known limitations rather than as truth.
What ROI should we expect from a CDP?
There is no reliable category figure. The return numbers quoted in the market typically come from vendor-commissioned economic models built from one or a small number of interviewed organisations, with attribution assumptions and risk adjustments applied, and those studies themselves tell readers to substitute their own estimates. Build an internal model: total cost against benefits you can actually evidence, counting nothing twice.
How long should a first CDP use case take before we judge it?
Long enough to log eligibility, delivery and exposure for a full business cycle and to compare the outcome against something. The trap is declaring success at the point the audience computes correctly, which measures the pipeline rather than the decision. Judge it on whether someone did something differently and whether that changed a result.
References
-
Twilio Segment, Identity Resolution Settings. Official product documentation. No publication or last-updated date shown on the page. Current live documentation. Vendor technical documentation: Twilio has a commercial interest in Segment, and the page establishes that identity rules must be configured. It does not establish that identity resolution is accurate, lawful or appropriate for any particular use case. https://segment.com/docs/personas/identity-resolution/identity-resolution-settings/
-
Information Commissioner’s Office, Data protection by design and by default. Regulatory guidance within the Guide to Accountability and Governance. No reference number. Publication date not shown; last updated 5 February 2026, when it was amended to reflect the Data (Use and Access) Act 2025. No under review notice shown on the page at the date of writing. UK regulator guidance, not legislation. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/accountability-and-governance/guide-to-accountability-and-governance/data-protection-by-design-and-by-default/
-
UK Parliament, Data (Use and Access) Act 2025, section 71, “The purpose limitation”. UK Public General Act, reference 2025 c. 18. Royal Assent 19 June 2025. UK primary legislation. https://www.legislation.gov.uk/ukpga/2025/18/section/71
-
Secretary of State, The Data (Use and Access) Act 2025 (Commencement No. 6) Regulations 2026. UK Statutory Instrument, reference SI 2026/82. Made 29 January 2026. Commenced section 71 on 5 February 2026. UK secondary legislation. https://www.legislation.gov.uk/uksi/2026/82/made
-
Information Commissioner’s Office, Respect people’s preferences. Section of the standalone Direct marketing guidance, not of the Guide to PECR. No section-specific publication or last-updated date shown; the parent Direct marketing guidance was published 5 December 2022 and last updated 28 April 2026. UK regulator guidance, not legislation. https://ico.org.uk/for-organisations/direct-marketing-and-privacy-and-electronic-communications/direct-marketing-guidance/respect-peoples-preferences/
-
Databricks, Building a complete and composable CDP on the Lakehouse. Vendor technical article. Published 17 October 2023. Current live article. Databricks sells the underlying data platform and has a commercial interest in warehouse-centred architecture. Describes the composable pattern; it does not independently establish lower cost, better outcomes or suitability for every organisation. https://www.databricks.com/blog/building-complete-and-composable-cdp-lakehouse
-
Bitol, under the LF AI and Data Foundation, Open Data Contract Standard. Open specification, version 3.1.0, approved 8 December 2025. Not an ISO International Standard, a W3C Recommendation or a statutory definition; contributors and participating tool vendors may have commercial interests in its adoption. A schema-valid contract can still contain weak thresholds or an incorrect business definition. https://bitol-io.github.io/open-data-contract-standard/v3.1.0/
-
Databricks, Guiding principles. Vendor architecture documentation. Updated 16 July 2026. Current. Databricks sells the platform described. Supports data-as-product, clear schema, lifecycle and semantic consistency as design principles; it is product-specific guidance and does not establish an industry-wide standard. https://docs.databricks.com/aws/en/lakehouse-architecture/guiding-principles
-
Forrester Consulting, commissioned by Twilio, The Total Economic Impact of Twilio Segment. Commissioned Total Economic Impact study, August 2024. Current hosted study. Commissioned by the vendor, which reviewed and provided feedback; Forrester states that it retained editorial control. The modelled 186 per cent three-year return, net present value and sub-six-month payback are based on one interviewed organisation together with attribution ratios, assumptions and risk adjustments over a three-year model. Not a universal expected return, not a competitive analysis, and the study directs readers to substitute their own estimates. https://tei.forrester.com/go/twilio/segment/?lang=en-us
Position dated 29 July 2026. Guidance, legislation and vendor documentation all change. References 1, 6, 8 and 9 come from parties with a commercial interest in the products they describe and are cited for their description of a pattern or an interface, not as independent evidence of value. UK data protection guidance continues to be revised following the Data (Use and Access) Act 2025. Verify any source before relying on it for a decision.
Use this guide as a source
If it settled an argument in your reporting, cite it, and add Herm as a preferred source so the next answer you get from search or a model comes from work with its definitions attached.