Conversational AI in customer service should be evaluated by what happens to the customer’s issue. Did the customer receive an accurate answer? Was an authorised change completed correctly? Did the system recognise when it could not help? Was the transfer to a person timely and useful?
A system has not succeeded merely because a conversation remained in an automated channel. It may have contained a contact by resolving the issue, but it may equally have caused abandonment, repeated contact or a complaint.
This article covers issue resolution, account servicing, agent assistance and escalation after a customer needs help. Product discovery, lead qualification and promotional commerce belong in conversational AI for marketing and commerce.
What This Article Delivers
- Resolution, not containment Why a conversation that stayed in an automated channel proves nothing, and what to measure instead.
- Ten service use cases Order status through complaint routing and service recovery, each with its authentication, permitted action and stop condition.
- Authentication by consequence Five action classes from public information to financial actions, with the assurance each one warrants.
- Escalation as part of resolution The nine conditions that should hand a customer to a person, and what the transfer must carry with it.
- An ROI model you build yourself Why there is no defensible universal saving per conversation, and why the denominator is a correctly resolved contact.
Marketing Conversation Versus Service Conversation
A service conversation begins when the customer needs to resolve, change or understand something relating to an existing relationship. Its primary outcome is correct and satisfactory resolution, not qualified demand, purchase or promotional engagement. For pre-purchase discovery, guided selling and lead qualification, see the conversational AI in marketing guide.
A service interaction may reveal a legitimate product need, but the existing issue should be resolved before any promotional suggestion is made. A customer who is distressed, vulnerable, disputing a charge or seeking a remedy should not be treated as an upselling opportunity.
Where This Article Sits in the Wider Cluster
For the wider organisational context, see AI and technology in marketing. Broader experience adaptation belongs in AI-driven personalisation, not in the core definition of support automation.
Governance principles are covered in the ethics of AI in marketing and ethical use of consumer data in marketing.
Service quality may influence retention, but churn causality and lifecycle programmes belong in reducing customer churn and personalised lifecycle engagement.
The Architecture of a Service Conversation
Task-oriented dialogue research distinguishes modular systems from more fully end-to-end ones. For customer service, separately governed components are often valuable, because the organisation needs to identify whether a failure occurred in intent recognition, knowledge retrieval, account access, action execution or escalation. (1)
| Component | Service function | Essential control |
|---|---|---|
| User interface | Receives questions and presents answers, controls, forms or voice prompts | Keep human and non-conversational alternatives visible |
| Speech recognition | Transcribes telephone or voice input | Confirm names, numbers, dates and action details, and test subgroup performance |
| Support-intent detection | Identifies order, billing, access, fault, complaint or other service needs | Route uncertain or multi-intent contacts for clarification or human review |
| Language model | Interprets and generates natural-language responses | It must not determine policy, entitlement or account authority by itself |
| Dialogue or orchestration layer | Tracks issue state and chooses retrieval, authentication, action or escalation | Encode mandatory steps and stop conditions |
| Retrieval | Finds policy, procedure, product and troubleshooting information | Use approved, versioned and effective sources |
| Tools and actions | Read orders, create cases, update permitted fields, process eligible returns or schedule support | Validate parameters and restrict authority by task |
| Customer identity | Associates the contact with the correct customer or account | Authenticate before exposing personal information or making protected changes |
| Permissions | Determines which records and actions are available | Use least privilege and step-up assurance for higher-impact actions |
| Business rules | Applies eligibility, refund, cancellation, complaint and service policies | Rules should come from authoritative systems rather than generated interpretation |
| Safety controls | Prevents unsupported advice, unsafe diagnostics and harmful action | Hard-stop medical, legal, financial and physical-safety risks |
| Logging | Records evidence, outputs, actions, authentication and escalation | Protect, minimise and retain records according to a documented purpose |
| Evaluation | Measures correctness, resolution and customer outcomes | Test complete journeys, not only model responses |
| Escalation | Transfers to support, complaints, fraud, safeguarding or vulnerability teams | Transfer context and avoid making the customer repeat the entire issue |
Retrieval-augmented generation can help produce answers based on external sources, but the original research evaluated selected benchmark tasks rather than customer-service policy accuracy. Retrieval can return the wrong document, an obsolete version or an incomplete passage. (2)
Different Service-System Types
| System type | Service use | Appropriate boundary |
|---|---|---|
| Scripted decision tree | Stable troubleshooting, status or eligibility sequences | Use when the permitted paths are limited and predictable |
| Intent-based natural-language system | Classifying common questions and selecting approved responses | Escalate ambiguous or overlapping intents |
| Retrieval-augmented generation | Explaining policy or troubleshooting from governed sources | Require source-quality, retrieval and answer testing |
| Tool-using assistant | Reading account state, creating cases or completing permitted actions | Authenticate and constrain every action |
| Voice interface | Telephone self-service or voice-enabled support | Treat transcription as a separate source of error |
| Fully automated action | Low-risk, reversible and clearly requested account operations | Do not automate discretionary remedies, unsafe instructions or ambiguous high-impact changes |
Research on machine-human handoff treats transfer as a prediction and collaboration problem connected to the state of the conversation and to likely satisfaction. Its experimental results are not a universal service benchmark, but they support the principle that escalation should respond to context and failure rather than to a fixed message count. (3)
Customer-Service Use Cases
Ten use cases are set out below. The first table covers what each conversation may do. The second covers when it must stop and how it is judged.
| Use case | Eligible issue | Authentication requirement | Permitted action | Customer alternative |
|---|---|---|---|---|
| Order status | Locate an existing order and explain its current fulfilment state | No account data before sufficient order or customer verification | Read status, carrier event and approved delivery information | Tracking page or support employee |
| Account and billing questions | Explain charges, balances, payment status or permitted account fields | Appropriate authentication, with step-up for sensitive changes or payment actions | Retrieve and explain authorised records, and update only approved low-risk fields | Billing or account specialist |
| Password or access recovery | Restore access through the approved recovery process | Secure proof appropriate to the account risk | Initiate or guide an authorised reset through protected channels | Secure recovery team or accessible alternative |
| Policy explanations | Explain current returns, warranty, cancellation or service policy | None for generic policy; authenticate before applying it to an account | Retrieve, cite and explain the current effective policy | Policy page or employee |
| Returns and cancellations | Check eligibility and complete a permitted return or cancellation | Authenticate before accessing the order or committing a change | Calculate the authorised option, show consequences and execute after confirmation | Self-service page or returns specialist |
| Fault diagnosis | Resolve a known product or service fault through approved steps | Depends on whether account, device or service data must be accessed | Run approved diagnostics, explain safe steps and schedule eligible repair | Technician, appointment or written troubleshooting route |
| Agent assist | Help an employee find evidence and prepare a response or action | Employee access plus the customer authentication required by the underlying task | Retrieve sources, suggest text, identify missing information and prepare permitted actions | Agent uses the ordinary knowledge and workflow systems |
| Conversation summarisation | Produce a draft case note or transfer summary | Authorised employee and system access to the contact | Summarise stated facts, completed checks, actions and unresolved issues | Manual case note |
| Complaint routing | Recognise dissatisfaction and create or route a formal complaint | Authenticate where the complaint concerns an account, while allowing accessible complaint initiation | Record the complaint, preserve evidence, identify urgency and send it to the correct queue | Direct complaint channel or complaint handler |
| Service recovery | Apply an approved remedy after a service failure | Authenticate and establish the relevant service event | Offer or apply only remedies within explicit authority | Human recovery specialist |
| Use case | Escalation condition | Evaluation metric | Harm guardrail | Stop condition |
|---|---|---|---|---|
| Order status | Missing order, conflicting status, material delay, suspected loss or requested remedy | Correct-status rate, repeat contact, delay escalation and disclosure errors | Do not reveal order or address information to an unverified person | Identity mismatch, unavailable source or contradictory tracking data |
| Account and billing questions | Dispute, suspected fraud, high-value consequence, policy exception or repeated failure | Correct resolution, repeat contact, complaint rate and unauthorised-action rate | Never invent a charge explanation or alter a record without authority | Failed authentication, unexplained discrepancy or request outside authority |
| Password or access recovery | Suspected takeover, inaccessible recovery method, repeated failures or vulnerable customer | Successful legitimate recovery, takeover incidents and repeat attempts | Never ask the customer to disclose an existing password, and protect recovery tokens | Attempt limit, suspected fraud or inability to establish identity |
| Policy explanations | Conflicting sources, legal interpretation, exception, complaint or material customer detriment | Source match, answer correctness, comprehension and repeat contact | Do not convert a general explanation into an unsupported decision about entitlement | No current authoritative source, or an unresolved conflict |
| Returns and cancellations | Exception, statutory-rights question, high charge, partial fulfilment or disputed condition | Correct completion, reversal, repeat contact, complaint and remedy accuracy | Show fees, deadlines and consequences before commitment, and do not obstruct cancellation | Uncertain eligibility, conflicting records or request outside delegated authority |
| Fault diagnosis | Safety risk, repeated failure, unknown fault, accessibility need or possible product defect | Correct resolution, first-contact resolution, repeat fault and unsafe-advice rate | Never generate hazardous electrical, medical or mechanical instructions | Any sign of danger, an unsupported procedure, or failure after the approved sequence |
| Agent assist | Policy conflict, missing evidence, vulnerability, complaint or high-impact decision | Source accuracy, agent correction, accepted assistance, resolution and handling effort | Show evidence and uncertainty, and never let a suggestion silently become a decision | No authoritative source, or an action beyond the employee's authority |
| Conversation summarisation | Sensitive ambiguity, multiple speakers, disputed facts or potential complaint evidence | Omission, attribution and fabrication rate, and correction effort | Distinguish customer statements from established facts and redact unnecessary sensitive data | Low confidence, inconsistent transcript or inability to attribute statements |
| Complaint routing | Vulnerability, safety, legal deadline, threatened harm, discrimination or repeated unresolved contact | Correct classification, timeliness, lost-complaint rate and repeat contact | Never downgrade, suppress or close a complaint to preserve automation metrics | Complaint intent, distress or any mandatory specialist condition |
| Service recovery | Discretionary remedy, high value, repeat harm, vulnerable customer or rejected offer | Correct remedy, recurrence, complaint outcome and acceptance without pressure | Do not require a customer to waive rights or accept promotion in exchange for help | Remedy outside authority, unresolved facts or customer rejection |
Authentication Should Match the Action
A generic policy explanation may need no identification. An order lookup, address change, cancellation or financial action may require progressively stronger assurance.
The NCSC’s customer-authentication guidance recommends selecting methods according to the organisation, the customer population, the security requirement and the usability implications. It is general online-service guidance rather than conversational-AI-specific guidance. (4)
| Action class | Example | Control |
|---|---|---|
| Public information | Generic opening hours or policy | No account authentication |
| Low-risk account reading | Order status with limited disclosure | Appropriate account or order verification |
| Account change | Contact preference or appointment change | Authenticated session and explicit confirmation |
| Financial or high-impact action | Payment, refund destination or material cancellation | Step-up authentication, action validation and audit record |
| Suspected fraud or coercion | Account takeover indicators or unusual instructions | Stop automated action and route to the relevant specialist |
The assistant should not collect passwords, full payment credentials or unnecessary identity documents in ordinary free text. Failed authentication must not leave the customer permanently unable to reach support.
Govern Support Knowledge as an Operational System
A language model does not make a knowledge base current. Customer-service knowledge needs owners, effective dates, approval states and withdrawal processes.
| Governance element | Requirement |
|---|---|
| Source owner | A named function responsible for accuracy and change approval |
| Effective date | The date from which a policy, procedure or message applies |
| Version history | An auditable record of what changed and why |
| Conflict handling | A rule that blocks or escalates when authoritative sources disagree |
| Retrieval coverage | Testing of whether customer wording finds the correct source |
| Answer evidence | The passages and version used for the response |
| Expiry control | Removal or de-prioritisation of obsolete information |
| Feedback loop | Structured correction from agents, complaints, quality review and failed contacts |
| Suspension | A way to disable a topic, tool or action when the knowledge cannot be trusted |
The ICO’s AI guidance covers accountability, transparency, fairness, security and data minimisation. It is under review following the Data (Use and Access) Act, so organisations should check current regulator material rather than treating an older implementation as permanently compliant. (5) (6)
Design Escalation as Part of Resolution
A customer should not have to discover the correct phrase for reaching a person. Escalation should be available when:
- the system cannot identify the issue;
- retrieval evidence is missing or contradictory;
- authentication fails through no fault of the customer;
- an authorised action fails;
- the customer repeats the same information without progress;
- the customer asks for a person;
- the interaction concerns a complaint, vulnerability, distress, fraud or safety;
- the system reaches a policy or authority boundary; or
- monitoring indicates that the system is not operating safely.
An effective escalation preserves the customer’s stated objective, verified identity state, completed checks, evidence retrieved, actions attempted, remaining uncertainty and unresolved questions. It must not conceal an error, convert an allegation into an established fact, or make the customer repeat information unnecessarily. A generated summary must therefore be reviewed or structured before it is passed on.
Escalation is also becoming a compliance question in the UK, not only a service-design one. The Data (Use and Access) Act amended the UK GDPR provisions on decisions taken solely by automated means that have a legal or similarly significant effect, and the ICO consulted between 31 March and 29 May 2026 on draft updated guidance interpreting those changes, with final guidance expected later in 2026. (7) A conversational system that refuses a refund, closes a complaint or ends a contract may be making that kind of decision, so the route to meaningful human review deserves specific legal advice rather than an assumption.
Customers in Vulnerable Circumstances
In UK financial services, FCA Finalised Guidance FG21/1 addresses firms’ treatment of vulnerable customers. It covers understanding customer needs, staff capability, practical action, communications, channel choice, monitoring and evaluation. The FCA page was updated in July 2026, while the underlying finalised guidance was published in February 2021. (8)
A service assistant should not infer a clinical or legal status from conversational style. It can, however, respond to operational signals such as:
- the customer saying they are distressed or unable to understand;
- repeated failure with the same step;
- bereavement, financial difficulty or coercion disclosed by the customer;
- speech, language or accessibility difficulties;
- imminent safety concerns; or
- a request for a trusted representative or an alternative channel.
The correct response may be slower service, a different format, a person with suitable authority, or suspension of an automated action.
Accessibility
WCAG 2.2 provides testable web-content criteria and should inform the interface, controls, errors, status messages and input alternatives. It does not by itself establish that a customer-service process is accessible or fair. (9)
Testing should include keyboard and assistive-technology use, cognitive load, timeouts, recovery from errors, plain-language comprehension, speech alternatives, and the accessibility of both authentication and human escalation.
Voice systems add speech-recognition risk. Research on one speech-to-text model found fabricated phrases in a small portion of the transcripts studied, and disproportionate effects in the aphasia data it examined. The study is model-specific and dataset-specific, but it shows why a service action must not rely on unconfirmed transcription of names, numbers, symptoms, complaints or financial instructions. (10)
Separate Service From Promotion
Service records and distress signals should not automatically become marketing opportunities. A genuine customer request for an additional product can be answered, but the system should first establish that:
- the service issue has been resolved or safely transferred;
- the customer initiated or clearly welcomed the commercial discussion;
- the recommendation is relevant to the stated need;
- the promotional purpose and material conditions are clear;
- vulnerability, distress or a complaint does not make promotion inappropriate; and
- declining the offer will not affect the customer’s remedy or access to service.
Detailed guided selling and lead-qualification conversations belong in the marketing-and-commerce article.
Measure Resolution, Not Containment Alone
| Metric | What it should establish |
|---|---|
| Task completion | Whether the intended service task was completed |
| Correct resolution | Whether the answer or action actually solved the eligible issue |
| Repeat contact | Whether the customer returned about the same issue within an appropriate period |
| First-contact resolution | Whether the issue was correctly resolved without avoidable further contact |
| Escalation quality | Whether escalation occurred when needed, and transferred useful context |
| Time to resolution | End-to-end elapsed time, including queues, retries and follow-up |
| Customer satisfaction | The customer's assessment, interpreted with response rate and sampling limits |
| Complaint rate | Formal and informal complaints attributable to the journey |
| Abandonment | Customers leaving before resolution, segmented by likely reason |
| Agent effort | Time and work required before, during and after an assisted contact |
| Action accuracy | Correct account, tool, parameters and confirmed outcome |
| Cost per correctly resolved contact | Total operating and repair cost divided by contacts meeting the resolution standard |
Containment should be reported only alongside correct resolution, repeat contact, abandonment and customer access to a person. A system that prevents transfer without solving the issue has not contained a successful contact.
Aggregate containment and cost measures can hide work transferred to frontline teams, repeat contacts and unresolved customer effort, so operational staff and structured transcript review should form part of evaluation.
Build an Organisation-Specific ROI Model
There is no defensible universal saving per automated conversation. Costs vary with the task, channel, authentication, integration, supplier contract, review burden and failure rate.
A useful service business case separates the following.
Benefits
- employee time genuinely avoided;
- repeat contact genuinely reduced;
- faster resolution where it benefits customers;
- fewer preventable errors;
- improved availability for tasks that can be completed safely;
- reduced training or search effort from effective agent assistance; and
- avoided downstream complaint or remediation work.
Costs
- platform and model usage;
- integration and security;
- knowledge ownership;
- testing and quality review;
- human escalation capacity;
- accessibility and language support;
- monitoring, incident handling and supplier oversight;
- incorrect actions and reversals;
- repeat contacts caused by failed automation; and
- complaints, customer harm and remediation.
A practical expression is net operating value equals evidenced benefits, less technology, governance, human-support, failure and remediation costs. The denominator for unit economics should be a correctly resolved eligible contact, not a message, a session or a nominally contained conversation.
Evaluate the Full Service Journey
NIST’s voluntary Generative AI Profile describes pre-deployment testing, red-teaming, representative evaluation, monitoring, incident response and deactivation planning. It is not UK law, but it provides a useful risk-management structure. (11)
| Evaluation method | Service application |
|---|---|
| Task-specific test sets | Representative status, billing, access, policy, return, fault and complaint scenarios |
| Retrieval accuracy | Correct current policy, procedure and account evidence |
| Answer correctness | Factual, procedurally correct and complete enough for the customer's task |
| Tool-execution accuracy | Correct account, permitted action, parameters and confirmation |
| Hallucination testing | Fabricated policy, entitlement, account state, remedy or troubleshooting step |
| Adversarial and misuse testing | Account takeover, prompt injection, abusive content, fraudulent refund requests and policy circumvention |
| Authentication testing | Bypass, session confusion, recovery weakness and inaccessible methods |
| Accessibility | Interface, language, authentication, timeout, error and escalation testing |
| Subgroup and language performance | Resolution, error, abandonment and escalation by supported language and relevant group |
| Escalation testing | Complaint, distress, vulnerability, uncertainty, repeated failure and human-request scenarios |
| Live controlled pilot | Correct resolution, repeat contact, customer outcome and operating cost against the existing route |
| Transcript review | Structured quality review with defined error categories |
| Post-deployment monitoring | Knowledge drift, tool incidents, complaint patterns, language shifts and supplier changes |
A model-level score can help compare components. It cannot establish that identity was verified, that the correct account was changed, that the customer understood the result, or that an escalation succeeded.
An Implementation Sequence Based on Risk
1. Select an Eligible Issue
Start with a frequent, sufficiently stable task for which correct resolution can be defined. Exclude discretionary, unsafe or weakly documented tasks.
2. Establish the Baseline
Measure current resolution, repeat contact, escalation, abandonment, effort, complaints and cost for the eligible population.
3. Define Authority and Stop Conditions
Specify what the assistant may read, explain, prepare and execute. Document the circumstances in which it must stop or transfer.
4. Govern Knowledge and Integrations
Assign owners to every source, and test the underlying tools independently of the conversational layer.
5. Test Offline
Use representative and adversarial scenarios, including authentication, accessibility, vulnerability and complaint cases.
6. Pilot With Constrained Exposure
Limit the audience, tasks and actions. Preserve the existing route and ensure that human capacity can absorb the escalation volume.
7. Compare Outcomes
Evaluate correctly resolved contact, repeat contact, complaints, abandonment, human effort and total cost. Do not scale because containment alone increased.
8. Expand Through a New Risk Decision
Each new action, language, channel or customer group changes the system. It should pass a new evidence gate rather than inherit approval automatically.
Legal and Ethical Operating Boundaries
A customer-service implementation should document:
- when and how the customer is told that the system is automated;
- which interactions require identity verification;
- how a customer reaches a person;
- how complaints, vulnerability and distress are recognised;
- which transcripts, recordings, summaries and action logs are retained;
- how service data is separated from promotional use;
- how unsupported financial, medical and legal advice is blocked;
- who can suspend a topic, an integration or the complete system; and
- how affected customers are identified and remediated after an incident.
Disclosure, consent or human handoff cannot compensate for an inaccurate policy engine, an insecure action, an inaccessible process, or an operating target that rewards unresolved containment.
The Design Principle
Conversational AI in customer service succeeds when it produces a correct and satisfactory resolution, or an effective escalation where automated resolution is not appropriate.
The most important question is not how many contacts the system contained. It is how many eligible customer needs were resolved correctly, safely and at an acceptable total cost.
References
-
Libo Qin, Wenbo Pan, Qiguang Chen, Lizi Liao, Zhou Yu, Yue Zhang, Wanxiang Che and Min Li, published by the Association for Computational Linguistics. End-to-end Task-oriented Dialogue: A Survey of Tasks, Methods, and Future Directions. Peer-reviewed conference paper, EMNLP 2023. DOI 10.18653/v1/2023.emnlp-main.363; ACL identifier 2023.emnlp-main.363. Published December 2023; no subsequent update. Not under review or in draft. Not UK law: a technical taxonomy rather than production assurance or an operating benchmark. https://aclanthology.org/2023.emnlp-main.363/
-
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel and Douwe Kiela. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Peer-reviewed conference paper, Advances in Neural Information Processing Systems 33 (NeurIPS 2020). Preprint identifier arXiv:2005.11401. Published 2020; no subsequent update. Not under review or in draft. Not UK law: most authors were affiliated with a commercial research laboratory with a technical interest in the method, and benchmark factuality improvements do not guarantee current policy retrieval or correct service resolution. https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html
-
Jiawei Liu, Kaisong Song, Yangyang Kang, Guoxiu He, Zhuoren Jiang, Changlong Sun, Wei Lu and Xiaozhong Liu, published by the Association for Computational Linguistics. A Role-Selected Sharing Network for Joint Machine-Human Chatting Handoff and Service Satisfaction Analysis. Peer-reviewed conference paper, EMNLP 2021. DOI 10.18653/v1/2021.emnlp-main.767; ACL identifier 2021.emnlp-main.767. Published November 2021; no subsequent update. Not under review or in draft. Not UK law: experimental results on two public datasets, which do not establish field performance, customer satisfaction or an optimal escalation policy. https://aclanthology.org/2021.emnlp-main.767/
-
National Cyber Security Centre (NCSC). Authentication methods: choosing the right type. UK government cyber-security guidance; not statutory. No reference number shown. Published 21 September 2022; status checked 28 July 2026. Not under review or in draft. UK source, but general customer-authentication guidance rather than conversational-AI-specific guidance, and not a substitute for sector requirements. https://www.ncsc.gov.uk/guidance/authentication-methods-choosing-the-right-type
-
Information Commissioner’s Office (ICO). Guidance on AI and data protection. UK regulator guidance; not a statutory code of practice. No reference number shown on the landing page. Updated 15 March 2023; status checked 28 July 2026. Under review: the live page states that, due to changes made by the Data (Use and Access) Act, the guidance is under review and may be subject to change. UK source and the applicable regulator for UK data protection. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/
-
Information Commissioner’s Office (ICO). How should we assess security and data minimisation in AI? Chapter of the UK regulator’s AI guidance. No separate reference number shown. Forms part of the AI guidance updated 15 March 2023; status checked 28 July 2026. Under review, along with the parent guidance, following the Data (Use and Access) Act. UK source, but it does not replace a complete security, authentication or data-protection assessment. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/how-should-we-assess-security-and-data-minimisation-in-ai/
-
Information Commissioner’s Office (ICO). ICO consultation on the draft guidance about automated decision-making, including profiling. Consultation on draft UK regulator guidance, updating the existing ADM and profiling guidance to reflect the Data (Use and Access) Act 2025. No reference number shown. Consultation opened 31 March 2026 and closed 29 May 2026; status checked 28 July 2026. DRAFT: this guidance was in draft and subject to consultation at the date of checking, with final guidance expected later in 2026, and a separate statutory code of practice on AI and automated decision-making is being prepared. UK source, but a draft carries no settled regulatory position and should not be relied on as a statement of current requirements. https://ico.org.uk/about-the-ico/ico-and-stakeholder-consultations/2026/03/ico-consultation-on-the-draft-guidance-about-automated-decision-making-including-profiling/
-
Financial Conduct Authority (FCA). FG21/1 Guidance for firms on the fair treatment of vulnerable customers. Finalised Guidance. Reference number FG21/1. Published 23 February 2021; webpage last updated 22 July 2026. Not under review or in draft. UK source, but sector-specific: non-financial organisations should not describe it as directly governing all their services. https://www.fca.org.uk/publications/finalised-guidance/guidance-firms-fair-treatment-vulnerable-customers
-
World Wide Web Consortium (W3C) Accessibility Guidelines Working Group, edited by Alastair Campbell, Chuck Adams, Rachael Bradley Montgomery, Michael Cooper and Andrew Kirkpatrick. Web Content Accessibility Guidelines (WCAG) 2.2. W3C Recommendation; an international technical standard. Reference WCAG 2.2. Published as a Recommendation 5 October 2023; the current Recommendation is dated 12 December 2024. Not under review or in draft: a stable Recommendation. Not UK law, although UK public-sector accessibility requirements refer to it. Web-content criteria do not measure every disability need or the accessibility of an entire support operation. https://www.w3.org/TR/WCAG22/
-
Allison Koenecke, Anna Seo Gyeong Choi, Katelyn X. Mei, Hilke Schellmann and Mona Sloane. Careless Whisper: Speech-to-Text Hallucination Harms. Peer-reviewed conference paper, ACM Conference on Fairness, Accountability, and Transparency (FAccT 2024). DOI 10.1145/3630106.3658996. Published June 2024; no subsequent update. Not under review or in draft. Not UK law: an independent academic study of one commercial speech-to-text model version and specific datasets, demonstrating a testing need rather than a universal error rate. https://doi.org/10.1145/3630106.3658996
-
National Institute of Standards and Technology (NIST), US Department of Commerce, authored by Chloe Autio, Reva Schwartz, Jesse Dunietz, Shomik Jain, Martin Stanley, Elham Tabassi, Patrick Hall and Kamie Roberts. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. Official government report and companion profile to the AI Risk Management Framework. Reference number NIST AI 600-1; DOI 10.6028/NIST.AI.600-1. Published 26 July 2024; publication page updated 8 April 2026. The parent AI RMF 1.0 is being revised, so the profile should be rechecked. Not UK law: voluntary, cross-sector US guidance and not a customer-service operating standard. https://doi.org/10.6028/NIST.AI.600-1
Position as at 28 July 2026. Regulatory guidance changes, and several sources here are unsettled. Reference 7 is a draft under consultation, references 5 and 6 are expressly under review following the Data (Use and Access) Act, and reference 11 sits under a framework being revised. All four should be rechecked, and the automated-decision position confirmed with a lawyer, before any decision is taken on the basis of this article.
Frequently Asked Questions
Why is containment a poor measure of customer-service AI?
Because it counts where a conversation ended rather than whether the customer's issue was dealt with. Containment rises when a system resolves an issue properly, but it also rises when a customer gives up, cannot find the route to a person, or is refused one. Those look identical in a containment figure and could not be more different for the customer. Report containment only next to correct resolution, repeat contact within an appropriate period, abandonment and whether access to a human remained effective. If containment improves while repeat contacts and complaints also rise, the system has moved work rather than removed it, and the saving is not real.
How much authentication should a support assistant require?
Match the assurance to the consequence of the action rather than applying one standard to everything. Explaining a general returns policy needs no identification at all. Reading an order status with limited disclosure needs appropriate order or customer verification. Changing a contact preference or an appointment needs an authenticated session and explicit confirmation. A payment, a refund destination or a material cancellation needs step-up authentication, validation of the action itself and an audit record. Where takeover indicators or unusual instructions appear, the system should stop and route to a specialist. Two rules apply throughout: never ask a customer to type an existing password or full payment credentials into a chat, and never let a failed authentication leave someone permanently unable to reach support.
When must a conversation be handed to a person?
Whenever the system cannot identify the issue, the retrieved evidence is missing or contradictory, authentication fails through no fault of the customer, an authorised action fails, or the customer repeats themselves without progress. Also whenever the customer simply asks, and whenever the interaction involves a complaint, distress, vulnerability, fraud or safety. The transfer itself matters as much as the trigger: it should carry the stated objective, the verified identity state, the checks already completed, the evidence retrieved and the questions still open, so the customer is not made to start again. Any generated summary needs reviewing or structuring so that uncertainty and customer allegations do not arrive as established fact.
How should we build a business case without a universal cost-per-conversation figure?
Build it from your own baseline, because the published per-conversation savings depend on task mix, channel, authentication burden, integration depth, contract terms and failure rate, none of which transfer. Measure current resolution, repeat contact, escalation, abandonment, effort, complaints and cost for one eligible issue first. Then count benefits conservatively as employee time genuinely avoided, repeat contact genuinely reduced, fewer preventable errors and less downstream remediation. Set against that the full cost, including model usage, integration, security, knowledge ownership, testing, escalation capacity, monitoring, incorrect actions and reversals. Divide by correctly resolved eligible contacts rather than sessions, because a contained conversation that solved nothing belongs in the cost column.
Use this guide as a source
If it settled an argument in your reporting, cite it, and add Herm as a preferred source so the next answer you get from search or a model comes from work with its definitions attached.