Artificial intelligence can recommend a product, rank an audience, predict customer behaviour, generate campaign material, or contribute to a decision about an individual. Those roles are not ethically equivalent.
A system suggesting alternative subject lines presents different risks from one suppressing offers for certain customers. A product recommender differs from a credit-eligibility system, even where both appear inside the same customer journey. A prediction is not a decision, but the business process acting on it can turn it into one.
Responsible AI in marketing therefore needs more than a set of principles. It needs controls attached to identifiable systems, named owners, documented boundaries, tests tied to foreseeable harms, meaningful intervention, and evidence that the controls still work after deployment.
This guide covers the governance of AI systems used in marketing. For broader organisational principles, see the ethical marketing framework. For the collection, profiling, sharing, retention and protection of customer information, see ethical use of consumer data in marketing.
Scope note: This is a governance framework, not legal advice. Legal classification depends on jurisdiction, intended purpose, deployment context and the consequences for affected people. Positions are stated as at 28 July 2026, and several of the dates below fall within days or months of that date.
What counts as AI in marketing
Governance should cover systems that infer, from their inputs, how to produce predictions, content, recommendations or decisions. That is broader than generative tools and customer-facing chatbots, and it matches the functional definition the EU AI Act uses (1).
Common marketing uses include:
- propensity and churn models;
- customer segmentation and lookalike modelling;
- product, content and offer recommendations;
- search, ranking and feed personalisation;
- dynamic creative optimisation;
- bid, budget and channel allocation;
- lead scoring;
- conversational systems;
- content, image, audio and video generation;
- sentiment and topic classification;
- identity resolution and audience matching;
- experimentation systems that adapt allocation automatically;
- tools summarising research, complaints or customer interactions;
- models supplied through marketing platforms, agencies or software vendors.
Do not restrict the inventory to things labelled AI. A conventional model, or a rules-plus-model system, can materially affect customers while the supplier describes it as analytics, optimisation, automation or decision intelligence. The reverse also holds: not every automated rule is an AI system, and legal classification should be assessed separately from internal governance scope.
I would put that more bluntly, having spent years selling and servicing this category. The deep AI and machine learning that genuinely moved client results was predictive segmentation. Clients chose a segment, the platform did the work behind it, and the measurable value showed up as return on ad spend once those segments were pushed into the advertising platforms. That is a real capability and it was worth paying for. It was also considerably more ordinary than the word AI implied to the people buying it.
Two things follow. When a vendor says AI, ask what the system actually infers and what changes for the customer as a result, because the answer is often a segment and a threshold. And when your own team says a tool is just analytics, apply the same question in reverse, because a segment and a threshold can decide who sees an offer.
Why AI creates distinct ethical risks
AI systems differ from ordinary campaign tools because their behaviour depends on training data, inferred patterns, optimisation objectives and deployment conditions that change.
Distinct risks include:
- errors occurring unevenly across groups;
- proxies reconstructing sensitive characteristics that were removed from the inputs;
- recommendations influencing choices without ever creating a formal decision;
- optimisation targets rewarding pressure, exclusion or repeated contact;
- models changing after a vendor update;
- generated content sounding confident while being false;
- unclear responsibility across a chain of model providers, platforms, agencies and deployers;
- explanations that describe the model rather than the customer consequence;
- human reviewers approving outputs routinely without independent judgement;
- feedback loops reinforcing popular products, established customers or historically successful segments;
- systems reused for purposes nobody assessed.
Governance therefore has to evaluate the whole sociotechnical system: model, data, interface, operating rule, business policy, human reviewer, customer route and commercial objective.
A technically accurate model can implement an unfair policy. Accurate customer data can still produce an unacceptable outcome. Removing a protected characteristic does not prove that related information is absent. An audit can find a problem without proving that discrimination has been eliminated.
Ethics, law and technical governance
Keep three questions apart.
Is the use lawful? This concerns applicable legislation and regulatory rules, including data protection, consumer protection, equality, advertising, intellectual property and sector requirements.
Is the system technically controlled? This concerns validity, reliability, security, data quality, testing, monitoring and change management.
Is the use ethically acceptable? This asks whether the objective, method and likely effect respect autonomy, treat affected groups fairly, and remain defensible even where the law permits them.
Compliance does not make a practice ethical. An organisation may be legally able to optimise an interaction and still decide that the pressure it puts on vulnerable customers is unacceptable. The reverse matters too: a well-intentioned project can breach a legal requirement.
Voluntary frameworks can help structure the work. The NIST AI Risk Management Framework is organised around govern, map, measure and manage (2), and ISO/IEC 42001 specifies requirements for an organisational AI management system (3). Neither replaces applicable law, a data protection impact assessment, or a decision about a specific use case. High-level intergovernmental principles serve a similar framing purpose without functioning as operational tests (4).
Build an inventory of AI use cases
An organisation cannot govern what it has not identified.
The inventory should reach systems built internally and those embedded in CRM and customer-data platforms, advertising platforms, email and journey tools, recommendation engines, agency workflows, analytics products, creative and copy tools, customer-service systems, browser extensions, foundation-model APIs, office software, and experimental prototypes running on live data.
Record the use case, not only the product name. One foundation model may support low-consequence copy ideation, a public chatbot, and a system summarising complaints for escalation. Those need different controls.
For each use case, document:
| Field | Question |
|---|---|
| Customer consequence | What changes for the customer? |
| Model role | Does the system advise, rank, generate or decide? |
| Data | What inputs and inferred attributes are used? |
| Objective | What is being optimised? |
| Error | What happens when the system is wrong? |
| Fairness | Which groups could experience unequal outcomes? |
| Explanation | What should the customer or the reviewer understand? |
| Human oversight | Who can intervene, and with what authority? |
| Monitoring | Which outcome and harm metrics are tracked? |
| Stop condition | What requires suspension or rollback? |
Add the system owner, supplier, model version, deployment date, jurisdictions, affected channels, dependency on personal data, approval status and reassessment date. An entry should stay open until the system is retired and its data, logs, credentials and downstream dependencies have been dealt with.
Expect the inventory to reveal a gap between capability and use, and treat that gap as a governance finding rather than an embarrassment. A meaningful share of the brands I worked with paid for sophisticated capability and used it to do something basic, which is like buying the best engine available and never leaving third gear. That sounds like a commercial waste story, and it is, but it has a governance edge. Unused capability sits in the estate unassessed, available to anybody who later decides to switch it on, and nobody reviews a feature that nobody is using. The inventory should record what a system is permitted to do as well as what it currently does, because the distance between those two is where unassessed uses appear.
Classify risk by decision and customer consequence
Do not classify by technical complexity. A simple model denying an important opportunity may need stronger governance than a complex model drafting internal copy.
Low consequence. Internal brainstorming, formatting or translation where no output is published without review, no personal data is entered, no customer is ranked or excluded, and errors are easy to detect. Controls: approved tools, confidentiality rules, factual review, content approval.
Moderate consequence. Product recommendations, send-time optimisation and campaign prioritisation, where the system changes customer exposure without determining eligibility. Controls should include documented objectives, data and model provenance, subgroup and accessibility testing, customer-experience monitoring, frequency and pressure limits, an opt-out or non-personalised route, and defined stop conditions.
High internal consequence. An organisational category, not a legal one. It can cover models allocating significant benefits or disadvantages, systems affecting vulnerable customers, decisions that are hard to reverse, extensive profiling, high-scale automation, sensitive inferences, and systems whose failure could cause substantial financial, reputational or dignitary harm. These need independent challenge, formal impact assessment, stronger evidence, senior approval and frequent monitoring.
Prohibited internally. Define what the organisation will not do even where the legal position is unsettled: fabricated testimonials, impersonating a real customer, employee or public figure, exploiting known distress, addiction or cognitive vulnerability, fabricating evidence for an advertising claim, suppressing complaint or cancellation routes, inferring highly sensitive traits purely to increase persuasive pressure, and releasing synthetic media without the review or disclosure its context requires.
EU AI Act classification
Ordinary product recommendations, content rankings and marketing personalisation are not automatically high-risk under the EU AI Act. Classification turns on intended purpose and whether the system falls within the listed categories.
AI used for targeted job advertisements or recruitment is listed in Annex III, as are systems assessing creditworthiness or determining credit scores, and certain systems used for life and health insurance risk assessment or pricing (1). A product recommender does not become high-risk because it uses profiling. An apparent recommendation that in practice determines access to credit or employment should be assessed by its real function rather than the team that owns it.
The timetable changed recently and is worth stating precisely. Following the Digital Omnibus on AI, published on 24 July 2026 and in force from 27 July 2026, the substantive Annex III high-risk obligations are scheduled to apply from 2 December 2027, and Annex I product-related high-risk obligations from 2 August 2028 (5).
Identify affected customers and stakeholders
Users are not the only people affected.
An impact map should consider customers who receive a recommendation, customers omitted from one, people represented in training or evaluation data, non-customers whose data may have entered a model, vulnerable people, children, people using assistive technologies, employees and agency staff expected to review outputs, creators and rights holders, businesses or sellers ranked by the system, customer-service teams, and anyone affected by a false generated statement or a synthetic likeness.
Absence is an outcome. Where a system consistently gives one group fewer opportunities to see an offer, the harm is underexposure rather than an incorrect prediction, and no accuracy metric will show it.
Consultation matters most where technical teams cannot determine on their own what customers would regard as an unacceptable consequence or an adequate explanation.
Data provenance and representativeness
Provenance means knowing where data came from, the conditions under which it was obtained, what was done to it, and whether its use remains appropriate for the current purpose.
Document source and collection period, original purpose, permission or lawful basis where personal data is involved, licences and restrictions, inclusion and exclusion criteria, the labelling process, missingness and known measurement error, geographic and temporal coverage, demographic or behavioural gaps, synthetic-data use, transformations and enrichment, lineage from source through to production features, and retention and deletion dependencies. Structured dataset documentation makes composition, collection and intended use visible rather than assumed (6).
Representativeness is use-case specific. A dataset can match the overall customer population and still be unsuitable for a particular language, channel, disability or low-frequency group. A model trained on customers who completed a purchase says little about those who abandoned because the earlier experience was inaccessible.
Diversity of records does not by itself remove bias. The labels, the objective, the collection process and the business rule can all encode inequity independently of the data mix.
Where personal data is used to develop or deploy a model, questions such as anonymity, lawful basis, and the consequences of unlawful processing at the development stage require case-by-case assessment (7). That is EEA guidance rather than UK law, but the underlying questions transfer. The full privacy lifecycle belongs in the organisation’s consumer-data governance rather than in model documentation. For collection, profiling, sharing, retention and customer rights, see the complete consumer-data lifecycle.
Define the optimisation objective
A model does not choose an ethical objective. People select the target, the constraints and the business process around it.
Maximise engagement can reward repeated interruption, emotionally provocative material, misleading curiosity gaps, content that is hard to disengage from, and disproportionate contact with highly responsive vulnerable users. Maximise conversion can under-serve customers who need more information before deciding.
Document the primary objective and the constraints together. A recommender might optimise relevant discovery while limiting contact frequency, repeated exposure, concentration on a narrow product set, unsuitable products, accessibility failures, exclusion disparities, pressure on vulnerable customers, and the use of sensitive or inferred attributes.
Review the policy before trying to improve the model. Fair model performance cannot repair a business rule that deliberately offers worse terms or reduced access to a group. The 4Ps ethics framework is the right place to test whether the underlying product, price, place or promotion policy is defensible. The AI assessment then tests whether the system implements that policy consistently and safely.
Bias, fairness and proxy discrimination
Bias can enter through problem framing, historical decisions used as labels, sampling, missing or mismeasured variables, the chosen target, feature engineering, proxies, model architecture, thresholds, ranking rules, interface design, human overrides, feedback loops, and the population the system is applied to. Bias is a lifecycle property rather than a training-data property (8).
Removing a protected characteristic is not enough. Postcode, device type, purchase history, language, channel preference and combinations of apparently neutral variables can all operate as proxies.
Fairness has to be defined before it can be measured. Candidate concerns include equal error rates, equal opportunity to receive a relevant offer, comparable recommendation quality, equitable exposure, avoidance of systematically worse prices or terms, accessibility, individual consistency, and protection against cumulative disadvantage.
These conflict. A system cannot satisfy every statistical definition at once, and no single metric settles the ethical question. One influential formal criterion defines fairness in terms of error behaviour across groups (9), while recommender-system research shows fairness remains context-dependent and involves competing user, provider and platform interests (10).
A fairness review should record the affected groups, why the selected fairness concept fits the use case, which other concepts were considered, acceptable disparity thresholds, confidence intervals and sample-size limits, the remediation decision, residual risks, and who approved what remains.
Bias cannot be declared permanently eliminated. The defensible claim is that specified risks were tested by stated methods and are being monitored.
Test performance across relevant groups
Aggregate performance conceals material differences.
Pre-deployment evaluation should test, where lawful and proportionate, error rates by relevant group, calibration, false-positive and false-negative rates, ranking position and exposure, recommendation coverage, refusal and fallback behaviour, accessibility, performance by language, region, device and channel, intersectional groups where samples permit, new against established customers, high-activity against low-activity customers, and behaviour under unusual seasonal or market conditions.
Protected-characteristic testing needs careful legal, privacy and security treatment. Do not collect sensitive data casually in the name of fairness. Establish that the testing is necessary, proportionate and permitted, then restrict access and retention.
Use several forms of testing. Offline evaluation on held-out and challenge datasets. Scenario testing against concrete journeys and foreseeable failure conditions. Red-teaming that deliberately attempts to produce harmful, misleading, discriminatory or policy-breaking behaviour. Shadow deployment, running the system without acting on its outputs. Controlled rollout with limited exposure, a defined evaluation period and a rollback route. Qualitative review of whether statistically acceptable outputs are contextually appropriate.
Before opening the data on any underperforming programme, it is worth knowing where the money usually is. Across a decade of client reviews, when results did not match expectations, the cause was almost always one of two things: somebody was reading the analytics incorrectly, or the test had been run without a proper setup. Neither is a modelling problem, and neither gets fixed by retraining. The same holds for fairness testing. A disparity that appears in a badly constructed test will waste weeks of remediation effort, and a real disparity hidden by a badly constructed test will ship. Validate the measurement before you trust what it says about the model.
A model can show equal prediction accuracy and still produce unequal outcomes, because different groups meet different thresholds, inventory, interfaces or review practices. Performance fairness and policy fairness have to be assessed separately.
Explainability and meaningful transparency
Transparency serves different audiences.
To regulators and internal reviewers, it may require technical documentation, data lineage, evaluation evidence, logs, model versions, supplier information, and the reasoning behind governance decisions.
To customers, it generally means understanding that AI materially influenced the interaction where that is relevant, what role it played, the principal information or behaviour affecting the outcome, the practical consequence, how to correct information, object or get help, and whether a meaningful alternative route exists.
A customer does not need the architecture. A model reviewer may need considerably more. UK guidance produced with an independent research institute sets out how to construct explanations of AI-assisted decisions for different audiences and purposes (11), and should be checked against post-DUAA updates before being relied on.
Explainability does not automatically create trust. A clear explanation can reveal that a system is unacceptable, and a plausible post-hoc explanation can misrepresent how the model actually behaved.
Documentation should distinguish system-level explanation from explanation of an individual output, global feature importance from the reason for a specific result, a model’s internal confidence from factual truth, understandable language from complete disclosure, and commercial confidentiality from information genuinely needed to challenge an outcome. Model cards are one established way of recording intended use, evaluation conditions, limitations and performance differences (12), and should be maintained as living records rather than written as promotional summaries.
Human oversight
A person somewhere in the workflow is not meaningful human oversight.
Oversight is meaningful only where the reviewer receives the information needed to evaluate the output, understands the intended use and limitations, has enough time, is trained to recognise the relevant errors and harms, can disagree without penalty, has authority to alter, pause or reverse the outcome, knows when escalation is mandatory, records the review where appropriate, and is not encouraged to approve routinely.
The process should state whether the person reviews before the output reaches the customer, examines samples afterwards, or responds only to complaints. Those are three different safeguards and should not be described interchangeably.
For significant solely automated decisions involving personal data, UK rules changed with the Data (Use and Access) Act 2025. Article 22 was replaced by Articles 22A to 22D, and the previous general restriction on significant solely automated decisions using ordinary personal data was relaxed, subject to lawful basis, transparency and safeguard requirements (13)(14). A right to human intervention arises where a system makes a decision about the person, the decision has a legal or similarly significant effect, and there was no meaningful human involvement before the outcome was applied. Safeguards include information about the decision, an opportunity to make representations, and a right to contest it (13). The detailed ICO guidance on this remains in draft: it was updated on 31 March 2026, consultation closed on 29 May 2026, and final guidance had not been issued at the time of writing (13).
A marketing prediction does not automatically fall inside those provisions. A lead, eligibility or pricing workflow may, once its output determines access, terms or opportunities.
Require review before action where the output could materially disadvantage an individual, the content makes a regulated or high-consequence claim, the system uses or infers sensitive information, an output refers to an identifiable person, generated content concerns health, finance, safety or legal rights, the result falls outside the validated operating range, confidence is low or evidence conflicts, the customer is vulnerable or has asked for human help, or the system proposes a use nobody approved.
The most common complaint I heard in a decade of quarterly business reviews was that the client’s team spent too much time operating the tool. It came up constantly and it never got solved, partly because saying it advertises how hard you are working. It matters here because that pressure is exactly what converts human oversight into a rubber stamp. If reviewing outputs is measured as time lost, reviewers will find ways to spend less of it, and the control will decay quietly while the process document still describes it as active. Test review quality directly through sampled decisions, disagreement analysis, reversal rates and reviewer feedback, and treat a falling override rate as a question rather than an achievement.
Recommendations, ranking and personalisation
A recommendation influences choice without determining eligibility, which is precisely why it escapes review.
Evaluate which products or messages can enter the candidate set, whether unsuitable items are excluded, whether commercial payments affect ranking, how popularity bias is controlled, whether customers can understand why something was recommended, whether repeated recommendations restrict discovery, how new and minority-interest items get exposure, whether customers can choose a non-personalised experience, whether the system exploits inferred urgency or vulnerability, and whether the design creates pressure through scarcity or repetition.
Recommendation fairness involves several parties at once, including customers, suppliers, creators and the platform, and the research cautions against assuming that a single abstract metric resolves what a fair recommendation means (10).
Measure more than clicks and conversion. Depending on the use case, monitoring may cover coverage and diversity, complaint and hide rates, repeat-exposure concentration, usefulness, unsuitable-item exposure, accessibility, subgroup performance, long-term customer outcomes, seller or creator exposure, and use of opt-out and alternative routes.
For privacy architecture and customer choice, see privacy-safe personalisation. For measurement design, see how to measure personalisation effectiveness.
Generative content and factual accuracy
Generative systems produce probable continuations, not verified facts. Fluency, specificity and model confidence are not evidence of truth, and confabulation and information integrity are recognised as risks requiring active management (15).
Controls should match the consequence of the content.
Low-consequence draft material. Use approved tools, keep confidential and unauthorised personal data out of prompts, retain human editorial responsibility, and review for accuracy, tone, representation and brand requirements.
Public factual material. Verify factual statements against authoritative sources, trace quotations to the original, confirm dates, names, prices, specifications and availability, require evidence for comparative or performance claims, check references rather than trusting generated citations, and record the responsible approver.
High-consequence material. For health, financial, safety, legal or eligibility-related content, restrict approved use cases, retrieve from controlled sources, prohibit unsourced generation, require suitably qualified review, test refusal and escalation behaviour, log source versions and approvals, and keep publication permissions narrow.
Prompt controls and blocked terms help without guaranteeing accuracy. Retrieval reduces some risks and can still surface outdated, irrelevant or misleading material.
Whatever the system produced, the campaign still has to pass ordinary advertising review. See the campaign substantiation and approval checklist for claims, misleading omissions, sponsorship and sign-off.
Synthetic media and disclosure
Content provenance asks how material was created. Content accuracy asks whether it is true. Neither proves the other. A machine-readable mark can show an image was generated without establishing that the depicted event occurred, and an unmarked photograph can mislead through selection, editing or context.
Disclosure decisions should weigh whether a reasonable person might believe a synthetic person or event is real, whether the content impersonates someone, whether a testimonial or endorsement is fabricated, whether an apparent customer interaction happened, whether the material concerns a public-interest matter, the medium involved, applicable legal and advertising requirements, and the prominence and timing of any disclosure.
There is no universal rule requiring every use of AI in advertising to be disclosed. UK advertising guidance asks whether omitting disclosure is likely to mislead, and disclosure cannot rescue an advertisement that is misleading in substance (16).
Under EU AI Act Article 50, duties differ by actor and content type: direct AI interaction generally requires notice unless it is obvious, providers of systems generating synthetic content must support machine-readable marking, deployers must label deepfakes, and certain AI-generated or manipulated text published to inform the public on matters of public interest must be labelled where the human-review and editorial-responsibility conditions are not met (1)(17).
Two dates matter and are frequently confused. The Article 50 obligations begin on 2 August 2026, days after this article’s stated position date. The transition running to 2 December 2026 is limited to Article 50(2) marking for systems already placed on the market before 2 August 2026, and is not a general grace period for the other transparency duties (18).
Avoid both extremes: disclosing every minor AI-assisted edit as though all uses were equivalent, and assuming nothing needs disclosure unless a statute uses the word AI.
Manipulation, vulnerability and customer autonomy
Optimisation becomes an ethical problem when it materially weakens a customer’s ability to choose voluntarily and knowingly.
Warning signs include modelling distress, impulsivity or financial pressure to raise conversion, repeated contact after signs of disengagement, time pressure the offer does not justify, personalised cancellation friction, withholding material information from customers predicted not to notice, using emotional inference to select pressure tactics, optimising for continued interaction with no customer-benefit constraint, and targeting children or vulnerable people with techniques they are less equipped to recognise.
The evidence here should be handled carefully in both directions. Research has found that matching messages to inferred psychological characteristics can affect behaviour in some experimental and field settings (19), and that work has been challenged on internal-validity grounds (20). The honest reading is that effects exist in some contexts, that their size and interpretation are contested, and that none of this establishes that all personalisation is manipulative.
The EU AI Act prohibits specified manipulative or deceptive techniques and the exploitation of vulnerabilities where the statutory conditions, including significant harm, are met (1). That threshold is not an ethical target. A practice can be unacceptable under an organisation’s own standards long before it reaches a legal prohibition.
Controls should include prohibited-targeting rules, vulnerability indicators used for protection rather than pressure, frequency and fatigue limits, cooling-off routes, simple opt-out and cancellation, non-personalised alternatives, monitoring for complaints and regret, and review of optimisation targets by legal, risk and customer teams. For customer transparency and choice, see privacy and personalisation.
Vendor and foundation-model governance
Buying an AI service does not transfer responsibility for how it is used.
Before approval, obtain and assess the provider identity and contracting entities, intended and prohibited uses, system documentation, training-data and copyright information available to downstream users, evaluation evidence relevant to your proposed use, known limitations, privacy and security terms, data-location and subcontractor information, whether prompts or outputs are used for training, retention and deletion arrangements, incident-notification commitments, update and deprecation policy, audit or assurance rights, continuity and exit arrangements, mechanisms for reporting harmful outputs, and the information you need to meet your own transparency duties.
Do not treat a vendor’s responsible-AI page as independent assurance. It states the vendor’s position. It does not establish that the system suits your population, your objective or your deployment process.
For foundation models, distinguish the model provider, any intermediate system provider, the deployer integrating the system, and the organisation publishing or acting on the output. EU general-purpose AI model obligations have applied since 2 August 2025, and the General-Purpose AI Code of Practice is a voluntary route intended to help providers demonstrate compliance with transparency, copyright and, for models with systemic risk, safety and security duties (21). A provider signing that code does not discharge the downstream deployer’s responsibility to assess its own use.
Contracts should require notification of material changes, new limitations, safety incidents and end-of-life dates. A silent model change invalidates your previous testing without telling you.
For protection of customer information and connected systems, see secure personalisation architecture and marketing data security.
Model changes and version control
Treat a change to a model, prompt, retrieval source, threshold or workflow as a controlled change wherever it may alter customer outcomes.
Record the model and API version, system instructions and prompt templates, retrieval sources and dates, safety settings, feature definitions, thresholds and ranking weights, approval date, evaluation dataset version, test results, known limitations, deployment scope and the rollback package.
Define which changes require no reassessment, limited regression testing, full impact reassessment, legal review, or renewed senior approval. Material triggers include a new foundation-model version, expansion to another country or language, use with a new customer group, a different optimisation objective, the introduction of sensitive or inferred data, movement from recommendation to automated action, removal of human review, large changes in error or complaint rates, a new customer consequence, or a change in legal classification.
Do not let a use case approved for internal drafting become a customer-facing decision tool through gradual scope expansion. That is the most common way an unassessed system reaches production.
Monitoring, drift and complaints
Pre-deployment testing describes a system under test conditions. Governance also needs evidence from real use.
Monitor four categories. Technical performance: accuracy and calibration, failure and refusal rates, latency and availability, out-of-distribution inputs, data and feature drift, model or vendor changes. Customer outcomes: exposure and exclusion, unsuitable recommendations, repeated contact, accessibility, completion, cancellation and correction routes, and long-term effects rather than immediate conversion. Fairness: subgroup performance, ranking and exposure disparities, threshold effects, override patterns, cumulative outcomes, and sample-size limitations. Harm and control effectiveness: complaints, contested outcomes, reviewer disagreement, reversals, generated inaccuracies, policy violations, privacy or security incidents, undisclosed synthetic media, and vendor incident notices.
A complaint is not only a customer-service issue. It can be evidence that the model, the policy, the explanation or the escalation route has failed. A complaint taxonomy should capture the system and version, the alleged harm, the affected group or accessibility need, whether the customer received an explanation, whether intervention changed the outcome, whether similar cases exist, whether monitoring failed to detect it, and whether the system should be suspended.
Do not read low complaint volume as safety. Customers may not know AI was involved, may not understand the outcome, and may find the complaint route inaccessible.
There is a related organisational failure I see constantly at the moment, and it is worth naming because it predicts how monitoring will actually be run. No brand I speak with tracks its own visibility inside AI assistants on its own initiative. The question only ever arrives from outside: a competitor appears to be winning at it, or somebody circulates a report with a league table in it. Underneath is a belief that being large, plus an SEO investment made a couple of years ago, will carry them. It will not. The behaviour matters here because monitoring an AI system you deployed has exactly the same shape. If nobody owns the number, nobody looks at the number, and the first time anyone examines the system will be after somebody external forces the question. Assign the metric to a person, put it on a schedule, and do not wait for a complaint or a competitor to start the review.
Incident response, rollback and retirement
An AI incident is broader than a security breach. It includes discriminatory or systematically unequal outcomes, widespread false claims, defamatory or harmful generated content, impersonation, exposure of confidential information, unsuitable recommendations, loss of required human review, unannounced vendor changes, incorrect eligibility or pricing, failure to honour opt-outs, material drift, and inability to reconstruct how an outcome happened.
The response plan should specify who can suspend the system, how customer impact is contained, which logs and evidence must be preserved, who assesses legal and regulatory notification, how affected customers are identified, when outcomes are reversed or re-reviewed, how corrections are communicated, whether similar systems are affected, who approves restart, and what evidence demonstrates remediation.
Rollback has to be technically possible and operationally rehearsed. A theoretical ability to return to an older model is worthless once the old data pipeline, API or approval process no longer exists.
Retirement should address removal of access and credentials, disabling automated jobs, retention or deletion of model inputs and logs, contractual termination, customer commitments, dependent systems, archive requirements, and lessons carried into successor systems.
AI marketing impact assessment
Complete an impact assessment before deploying a material new use case, and repeat it after significant change. Where personal data is involved, this sits alongside rather than instead of a data protection impact assessment, and large-scale profiling and marketing data matching both appear on the ICO’s high-risk list (22).
Purpose and necessity. What problem is being solved? Why is AI needed? What less intrusive alternative was considered? What customer benefit is expected? What commercial objective could conflict with it?
System and data. What system, model and supplier are involved? What data is used or inferred? What is its provenance? Is personal or sensitive information involved? What limitations are known?
Consequences and affected people. Who is affected directly and indirectly? What changes for them? What happens when the system is wrong? Can the result accumulate? Are vulnerable people or children affected?
Fairness. Which groups might experience different outcomes? Which fairness concept applies? Which metrics and qualitative tests will be used? What disparity triggers remediation or suspension? How are small samples handled?
Transparency and autonomy. What must internal reviewers know? What must customers know? Is disclosure required? Can customers correct information, object or choose an alternative? Could the design impair autonomy?
Human oversight. Which decisions require review? Who reviews them? What information, training and authority do they have? How are overrides recorded? Who can stop the system?
Vendor and operational control. What documentation and contractual rights exist? How will updates be notified? Can the organisation monitor the relevant outcomes? Is rollback possible? What happens if the vendor withdraws the service?
Approval and residual risk. Which controls are required before launch? Who accepts the residual risk? What is the monitoring period? What are the stop conditions? When is reassessment due?
Lifecycle guidance for AI system impact assessments exists as an international standard (23), alongside the organisation-level management-system structure noted earlier (3). Those support consistency. They do not determine whether a specific marketing use is lawful or acceptable.
Governance responsibility matrix
| Activity | Marketing owner | Data and AI team | Legal and privacy | Risk and compliance | Content or CX | Procurement and security | Senior approver |
|---|---|---|---|---|---|---|---|
| Define use and objective | Accountable | Consulted | Consulted | Consulted | Consulted | Informed | Informed |
| Inventory entry | Accountable | Responsible | Consulted | Informed | Informed | Consulted | Informed |
| Data and model documentation | Consulted | Accountable | Consulted | Informed | Informed | Consulted | Informed |
| Legal classification | Consulted | Consulted | Accountable | Consulted | Informed | Informed | Informed |
| Fairness-test design | Consulted | Responsible | Consulted | Accountable challenge | Consulted | Informed | Informed |
| Content and authenticity review | Accountable | Consulted | Consulted | Informed | Responsible | Informed | Informed |
| Vendor assessment | Consulted | Consulted | Consulted | Consulted | Informed | Accountable | Informed |
| Deployment approval | Responsible | Responsible | Approval within remit | Approval within remit | Approval within remit | Approval within remit | Accountable for higher-risk uses |
| Monitoring | Accountable | Responsible | Consulted | Independent challenge | Responsible for customer signals | Consulted | Receives material reports |
| Incident suspension | Authorised | Authorised | Authorised within remit | Authorised | Escalates | Authorised for security events | Accountable for restart |
| Retirement | Accountable | Responsible | Consulted | Informed | Informed | Responsible for vendor access | Informed |
Titles matter less than explicit authority. Name the roles and their alternates.
AI marketing checklist
Inventory and boundaries
- The use case is recorded in the AI inventory.
- The model’s role, whether it advises, ranks, generates or decides, is documented.
- Approved and prohibited uses are explicit.
- Customer consequences and affected groups are mapped.
- Legal and internal risk classifications are recorded separately.
Data and documentation
- Data and model provenance are documented.
- Training, evaluation and production data are distinguished.
- Representativeness and known gaps are recorded.
- Documentation states intended uses, limitations and evaluation conditions.
- Personal-data requirements run through the organisation’s privacy process.
Testing
- Pre-deployment evaluation covers realistic operating conditions.
- Relevant subgroup performance has been tested lawfully and proportionately.
- Proxy risks have been considered.
- Red-teaming covers harmful, misleading and policy-breaking outputs.
- Accessibility has been tested.
- Human-review effectiveness has been tested.
- Residual limitations are visible to operators.
Generative content
- Prompt and output controls match the content’s consequence.
- Public factual claims are independently verified.
- Quotations and citations are traced to originals.
- Intellectual-property and provenance checks are defined.
- Synthetic-media disclosure has been assessed.
- A named person retains editorial responsibility.
Human oversight and customer routes
- Reviewers have time, knowledge and authority to intervene.
- Mandatory escalation conditions are explicit.
- Customers can reach an effective human route where needed.
- Opt-out or non-personalised alternatives exist where appropriate.
- Corrections, objections and complaints reach the system owner.
Vendors
- Vendor documentation has been reviewed rather than accepted at face value.
- Contractual use restrictions match the intended use.
- Prompt, output, training and retention terms are understood.
- Update and incident notifications are required.
- Audit or assurance evidence is available.
- Exit, continuity and rollback arrangements exist.
Monitoring and incidents
- Technical, outcome, fairness and harm metrics are defined and owned.
- Thresholds trigger investigation, suspension or rollback.
- Complaints are linked to system and version records.
- Incident logs preserve decisions and evidence.
- Reassessment dates are scheduled.
- Retirement procedures cover data, access, contracts and dependencies.
Frequently Asked Questions
What are the main ethical risks of AI in marketing?
Unfair outcomes across groups, opaque personalisation, unsuitable recommendations, manipulative optimisation, factual errors, fabricated or misleading content, inadequate human authority, unauthorised data use, weak vendor governance, and failures that continue undetected after deployment because nobody owns the monitoring.
Is ordinary marketing personalisation high-risk under the EU AI Act?
Not automatically. Ordinary product personalisation is not expressly listed as an Annex III high-risk category. Classification depends on intended purpose and actual function, so a system used for recruitment, creditworthiness, credit scoring or specified insurance decisions may fall within a listed category even when it sits inside a marketing or customer-acquisition journey.
Are AI-generated advertisements always required to be disclosed?
No universal rule requires disclosure of every AI-assisted advertisement. It depends on whether omission would mislead, and on specific duties covering interaction notices, machine-readable marking, deepfakes and certain public-interest text. The underlying advertisement must be truthful and non-misleading regardless, and disclosure cannot rescue one that is not.
Does removing protected attributes make a model fair?
No. Other variables can act as proxies, historical labels may encode unequal treatment, and the business objective itself may be unfair. Fairness requires context-specific analysis of the data, model performance, thresholds, exposure and the underlying policy.
Does a fairness audit guarantee non-discrimination?
No. An audit reflects its data, methods, definitions and time period. It can miss small groups, intersectional effects, future drift, and harms its metrics were never designed to capture. The defensible claim is that specified risks were tested by stated methods and are being monitored.
What makes human oversight meaningful?
The reviewer must understand the system, receive the relevant information, have enough time, and hold genuine authority to disagree, change the outcome, escalate or stop the process. A routine approval click is not oversight, and a falling override rate should be treated as a question rather than an achievement.
How should algorithmic bias be tested?
Start by defining the potential harm and the fairness concept that fits it. Test error, calibration, exposure, ranking and customer outcomes across relevant groups, then combine statistical testing with scenario review, red-teaming, accessibility testing and post-deployment monitoring. Validate the test setup before trusting what it says about the model.
How should foundation-model vendors be assessed?
Assess documentation, the data and copyright information available to deployers, evaluation evidence, privacy and security terms, retention, permitted use, updates, incidents, audit rights, subcontractors, deprecation and exit arrangements. Then test the complete downstream application in its actual marketing context, because vendor assurance describes the model rather than your deployment.
Is explainable AI always more trustworthy?
No. An explanation can reveal a harmful practice, omit the relevant cause, or create unjustified confidence. Explanations need to be accurate enough for their purpose and connected to a real route for challenge or intervention.
What is the difference between content provenance and accuracy?
Provenance records how content was created or altered. Accuracy concerns whether it is true. Synthetic-content marking does not verify a claim, and human-created content is not automatically accurate.
Conclusion
Responsible AI marketing starts with a specific system and a specific consequence, not with a policy statement.
A team should be able to say what the system does, what it is permitted to do, whose interests are affected, what is being optimised, how performance differs across groups, who can intervene, and what evidence would cause the system to be stopped.
That requires an inventory, defensible risk classification, documented data and models, fairness and performance testing, explanations customers can act on, real human authority, controlled generative workflows, vendor accountability, version control, monitoring, complaints and incident handling, and rehearsed rollback and retirement.
AI does not take on moral responsibility for a marketing outcome. Responsibility stays with the organisations and the people who choose the objective, approve the system and act on its outputs. For the organisation-wide principles surrounding those choices, return to the ethical marketing framework.
References
- European Parliament and Council of the European Union, Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), EU regulation, adopted 13 June 2024, published 12 July 2024, ELI http://data.europa.eu/eli/reg/2024/1689/oj. EU law, not UK law. https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
- Elham Tabassi, National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), ref. NIST AI 100-1, published 26 January 2023, DOI 10.6028/NIST.AI.100-1. Voluntary framework; not certification, a legal safe harbour, or UK law. https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-ai-rmf-10
- International Organization for Standardization and International Electrotechnical Commission, ISO/IEC 42001:2023, Information technology, Artificial intelligence, Management system, international standard, 2023. Voluntary standard; ISO sells access to the full text, and only the official overview was reviewed here. https://www.iso.org/standard/42001
- Organisation for Economic Co-operation and Development, OECD AI Principles overview, intergovernmental principles, adopted May 2019, updated May 2024. Non-binding unless incorporated into national law. https://oecd.ai/en/ai-principles
- European Parliament and Council of the European Union, Regulation (EU) 2026/1744 amending Regulations (EU) 2024/1689, (EU) 2018/1139 and (EU) 2023/1230 as regards the simplification of the implementation of harmonised rules on artificial intelligence (Digital Omnibus on AI), EU regulation, adopted 8 July 2026, published 24 July 2026, in force 27 July 2026. https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=OJ:L_202601744
- Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daume III and Kate Crawford, Datasheets for Datasets, Communications of the ACM, 64(12), 2021, DOI 10.1145/3458723. Proposes documentation practice rather than a legal standard. https://dl.acm.org/doi/10.1145/3458723
- European Data Protection Board, Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models, Article 64 opinion, adopted 18 December 2024. EEA guidance, not UK law. https://www.edpb.europa.eu/documents/opinion-of-the-board-art-64/opinion-282024-on-certain-data-protection-aspects-related-to_en
- International Organization for Standardization and International Electrotechnical Commission, ISO/IEC TR 24027:2021, Information technology, Artificial intelligence (AI), Bias in AI systems and AI aided decision making, technical report, published 5 November 2021. Voluntary technical report; the public abstract was reviewed. https://www.iso.org/standard/77607.html
- Moritz Hardt, Eric Price and Nati Srebro, Equality of Opportunity in Supervised Learning, Advances in Neural Information Processing Systems 29, 2016. One formal fairness criterion among several; it does not establish a universal ethical or legal definition. https://proceedings.neurips.cc/paper_files/paper/2016/hash/6a9659feb1216f14f7384ba499518b38-Abstract.html
- Yashar Deldjoo, Dietmar Jannach, Alejandro Bellogin, Alessandro Difonzo and Dario Zanzonelli, Fairness in recommender systems: research landscape and future directions, User Modeling and User-Adapted Interaction, 34, 59 to 108, published online 24 April 2023, DOI 10.1007/s11257-023-09364-z. An academic survey describing a developing field rather than settled practice. https://link.springer.com/article/10.1007/s11257-023-09364-z
- Information Commissioner’s Office and The Alan Turing Institute, Explaining decisions made with AI, regulatory guidance produced with an independent research institute. Should be checked against post-DUAA updates. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/explaining-decisions-made-with-artificial-intelligence/
- Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji and Timnit Gebru, Model Cards for Model Reporting, conference paper, FAT* 2019, DOI 10.1145/3287560.3287596. A documentation proposal, not independent certification; several authors held industry affiliations at publication. https://dl.acm.org/doi/10.1145/3287560.3287596
- Information Commissioner’s Office, Automated decision-making, including profiling, draft regulatory guidance, updated 31 March 2026. Consultation closed 29 May 2026 and final post-consultation guidance had not been issued at the time of writing. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/individual-rights/automated-decision-making/
- Information Commissioner’s Office, The Data Use and Access Act 2025 (DUAA): what does it mean for organisations?, regulatory guidance, published 19 June 2025, materially updated 19 June 2026. https://ico.org.uk/about-the-ico/what-we-do/legislation-we-cover/data-use-and-access-act-2025/the-data-use-and-access-act-2025-what-does-it-mean-for-organisations/
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, ref. NIST AI 600-1, published 26 July 2024, updated 8 April 2026, DOI 10.6028/NIST.AI.600-1. Voluntary profile; identifies risks including confabulation and information integrity without quantifying risk for any specific application. https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence
- Committee of Advertising Practice, Disclosure of AI in Advertising: Striking the Balance Between Creativity and Responsibility, CAP News guidance, published 29 May 2025. Advertising self-regulatory guidance. https://www.asa.org.uk/news/disclosure-of-ai-in-advertising-striking-the-balance-between-creativity-and-responsibility.html
- European Commission, Directorate-General for Communications Networks, Content and Technology, Guidelines on transparency obligations for providers and deployers of AI systems, official guidance, 20 July 2026. https://digital-strategy.ec.europa.eu/en/library/guidelines-transparency-obligations-providers-and-deployers-ai-systems
- European Commission, Directorate-General for Communications Networks, Content and Technology, Transparency obligations under Article 50 of the AI Act, official FAQ, last updated 24 July 2026. https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act
- Sandra C. Matz, Michal Kosinski, Gideon Nave and David J. Stillwell, Psychological targeting as an effective approach to digital mass persuasion, Proceedings of the National Academy of Sciences, 114(48), 2017, DOI 10.1073/pnas.1710966114. Rests on specific experiments and platform conditions; not proof of universal effectiveness or harm. https://www.pnas.org/doi/10.1073/pnas.1710966114
- Dean Eckles, Brett R. Gordon and Garrett A. Johnson, Field studies of psychologically targeted ads face threats to internal validity, Proceedings of the National Academy of Sciences, 2018, DOI 10.1073/pnas.1805363115. A methodological critique of the interpretation of those field studies; it does not establish that psychological targeting has no effect. https://www.pnas.org/doi/10.1073/pnas.1805363115
- European Commission and European AI Office, The General-Purpose AI Code of Practice, voluntary code and official explanatory page, published 10 July 2025. A voluntary compliance route developed through a Commission-facilitated multistakeholder process; signatories include commercial model providers. https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai
- Information Commissioner’s Office, Examples of processing ‘likely to result in high risk’, Article 35(4) list and accompanying guidance, no publication or update date displayed. Under review following the Data (Use and Access) Act 2025. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/accountability-and-governance/data-protection-impact-assessments-dpias/examples-of-processing-likely-to-result-in-high-risk/
- International Organization for Standardization and International Electrotechnical Commission, ISO/IEC 42005:2025, Information technology, Artificial intelligence (AI), AI system impact assessment, international standard, published 28 May 2025. Voluntary standard; conclusions here rest on the public official abstract. https://www.iso.org/standard/42005
Position stated as at 28 July 2026. Regulatory guidance changes, and several of the sources above are marked by their publishers as under review. Check the linked sources before relying on any statement of the current legal position.
Use this guide as a source
If it settled an argument in your reporting, cite it, and add Herm as a preferred source so the next answer you get from search or a model comes from work with its definitions attached.