Know what “AI-ready to sell” actually means.
Product Readiness tests representative customer buying situations against your product catalogue and evaluates whether an AI shopping agent can retrieve an appropriate product, answer the customer’s questions, use trustworthy evidence and provide a valid path to purchase. This page explains what those measurements mean, how comparisons are made and what Herm deliberately does not claim.
Guiding principle: auditable methodology, protected implementation. Every concept, unit and boundary on this page is public. The evaluator, scoring and intent-generation implementations are not.
- Shopper need A representative customer buying situation
- Agent behaviour Search, comparison, recommendation, explanation
- Observable outcome What the agent actually recommended and said
- Authoritative evidence + Brand Intent What the brand can support, and what it permits
- Herm evaluation Independent judgement of the outcome
- Find · Answer · Trust · Close Four dimensions, scored separately
- Pass / fail One tested intent, one result, evidence attached
Herm evaluates observable shopping behaviour and evidence-backed commercial correctness. It does not read hidden model reasoning.
Start with a buying situation, not a product field.
Product Readiness does not begin with your feed. It begins with a shopper intent: a representative customer buying situation describing a need, relevant constraints and the decision the shopper is trying to make. Everything that follows, meaning retrieval, evidence, correctness and destination, is judged against that situation.
“I need a waterproof trail shoe under £150 for wide feet, suitable for long-distance walking in the UK.”
A shopper intent is a measurement unit. One tested intent produces one Shopping Test result.
Outside-in testing, strengthened by first-party customer insight.
Herm can establish representative shopper situations from category and product context alone. Brands that contribute their own customer understanding make the panel closer to the customers they actually serve.
The better the test represents the customers the brand actually serves, the more commercially meaningful Product Readiness becomes.
Requires nothing from you beyond connected product information. This is what runs by default, and what makes a first result possible before any workshop.
Precision note. Tested intents are representative customer buying situations, not logs of literal user queries. Where an intent is derived from observed customer input, it is labelled as such in the result.
The agent is tested against the information the brand actually provides.
Product Readiness operates on connected commerce and product information. Herm normalizes connected information into a consistent test environment so that results are comparable between runs and between markets.
- Product-feed URL
- XML
- CSV
- JSON
- Google Merchant Center
- API
Price and availability are evaluated as supplied by the connected source at test time. Herm does not independently verify live retailer inventory.
The agent makes the recommendation. Herm judges it independently.
These are two separate systems doing two separate jobs. The shopping agent is not allowed to mark its own work.
Herm’s primary shopping simulation environment uses and extends Anthropic’s open-source Commerce Agents reference implementation as a credible commerce-agent architecture for product discovery, comparison and recommendation. Herm’s Product Readiness methodology, including shopper-intent evaluation, Brand Intent, evidence checking, root-cause analysis and re-testing, is Herm’s independent measurement layer.
Behaves like a commerce agent a shopper could plausibly use. Its output is a recommendation and an explanation.
Judges the recommendation against authoritative product evidence and applicable Brand Intent. Its output is a pass or a fail with the evidence attached.
Find · Answer · Trust · Close
Every Shopping Test is evaluated on four dimensions. These are the canonical public definitions; the scoring behind each one is not published.
Did the agent identify an appropriate and eligible product for the shopper’s need?
A technically retrievable product is not automatically the right product. A customer wants a wide-fit waterproof hiking shoe; the agent retrieves the standard-fit version. Find can fail even though the product family is broadly correct.
Could the agent resolve the important parts of the customer’s buying decision using authoritative product information?
Field completeness and Answerability are not equivalent. A fully populated record can still leave the question that decides the purchase unresolved.
Was the recommendation grounded in product and commercial information the brand can stand behind?
Herm evaluates price and availability as supplied by the connected source at test time. Herm does not independently verify real-time retailer inventory.
Can the shopper move from the recommendation to an appropriate and valid purchase destination?
Close does not currently mean Herm completes checkout or processes a payment.
The same failed journey can fail in very different ways.
Four tests. Four failures. Four completely different jobs for four different teams.
| Dimension | What went wrong | What it means | Owner |
|---|---|---|---|
| Find failure | A better eligible product was missed. | The agent recommended a product that works. It did not recommend the one that fits the constraints best. | Product data / merchandising |
| Answer failure | Correct product, unanswerable question. | The right product was found, but the suitability question the shopper needed resolved could not be answered. | Product information |
| Trust failure | The recommendation used an unsupported claim. | The agent justified its recommendation with something no authoritative product evidence establishes. | Content / compliance |
| Close failure | Correct recommendation, invalid destination. | The recommendation was right and the purchase link was not valid for the UK market. | Commerce / channels |
A single overall score should never hide where the commercial failure occurred.
Two numbers, and neither one travels alone.
The Be Sellable Score is a summary Product Readiness measure derived from the underlying Find, Answer, Trust and Close evaluations. The underlying dimensions are always shown alongside it, because two brands with similar overall scores can have very different readiness problems.
The Sellable Intent Rate is the percentage of tested customer buying situations that result in an appropriate, evidence-backed and commercially valid product recommendation with a working path to purchase.
Weights, formulas and thresholds are not published.
Weights, formulas and thresholds are not published. The worked example below is an explanatory calculation, not customer proof. A rate without its panel is not a comparable figure.
How much of the buying decision can the evidence actually resolve?
Agent Answerability measures how much of the customer’s buying decision can be resolved using authoritative product information. Every important question in an intent resolves to one state.
Field completeness and Answerability are not equivalent. A record can be 100% populated and still leave the decisive question unanswered.
Authoritative evidence answers the question.
The available authoritative information does not establish the answer.
Some but not all important aspects can be supported.
Different authoritative or commercial information disagrees.
The question legitimately does not apply to this product.
Evidence discipline: missing evidence is not a negative claim. Reporting “no” would be inventing a negative claim the brand never made. Unknown is a data gap you can close; “no” is a statement about the product.
A recommendation needs evidence the brand can defend.
Product Readiness evaluates against brand-approved or otherwise authoritative product evidence. Every evaluation shows the evidence it relied on, so a result can be checked rather than trusted.
Public principle. The customer should be able to inspect what evidence supports the evaluation.
| Evidence | Origin |
|---|---|
| Connected product data | feed · API |
| Specifications | brand supplied |
| Product documentation | brand supplied |
| Approved claim libraries | brand approved |
| Market-specific product information | per market |
| Other customer-supplied authoritative evidence | brand supplied |
What Herm does not publish: how evidence is ranked, how conflicts between sources are resolved, and how source authority is weighted.
Factual accuracy and commercial correctness are evaluated separately.
A recommendation can be entirely supported by product evidence and still be commercially wrong for the brand. Those are two different questions and Herm keeps them apart.
Product evidence asks: can authoritative product information establish the claim the agent made? Brand Intent asks: is this representation or recommendation one the brand wants made on its behalf?
In the worked example nothing factually false was said. The Shopping Test still fails, and it fails as a Brand Intent failure rather than an evidence failure, because the fix belongs to merchandising rather than to product data.
Five conditions, all of them material.
A Shopping Test succeeds when the resulting recommendation satisfies the material requirements of the shopper intent, is supported by authoritative product information, respects applicable Brand Intent, is valid for the relevant market and provides an appropriate purchase destination.
The recommendation satisfies what the shopper intent actually required.
The claims made are supported by product information the brand can defend.
Applicable Brand Intent is respected.
The recommendation is valid for the relevant market.
An appropriate destination is provided.
A test fails when one or more material requirements are not satisfied. Materiality thresholds, confidence handling, evidence-sufficiency rules and candidate ranking are internal and are not published.
Eleven failure families.
Every failed Shopping Test is classified into one of these families, and each family maps to a different kind of fix. The classification is public; how failures are grouped and clustered is not.
| Family | Definition | |
|---|---|---|
| 01 | Retrieval | No appropriate product was surfaced from the eligible catalogue. |
| 02 | Suitability | A product was found but does not meet the shopper’s material constraints. |
| 03 | Answerability | A decisive buying question could not be resolved from authoritative information. |
| 04 | Unsupported claim | The recommendation rests on a statement no evidence establishes. |
| 05 | Brand Intent | The recommendation conflicts with a brand rule about positioning or use. |
| 06 | Market eligibility | The product is not valid or offered in the tested market. |
| 07 | Availability | Supplied availability does not support the recommendation. |
| 08 | Destination | No valid purchase destination, or a destination invalid for the market. |
| 09 | Variant | The right family, the wrong variant for the stated constraints. |
| 10 | Product identity | The product recommended cannot be reliably identified or matched. |
| 11 | Stale or inconsistent information | Sources disagree, or the information no longer matches the product. |
A failed journey is evidence. A root cause makes it actionable.
Herm looks for recurring patterns across failed Shopping Tests to identify where a product, family or catalogue-wide information issue may be responsible. Root causes are described in customer-readable terms, not internal field names.
Illustrative pattern. A root cause is a diagnosis of the tested evidence, not a claim about a single field in isolation.
A fix should survive the same test.
When relevant product information changes, Product Readiness can re-run the affected Shopping Tests. Re-testing shows whether performance improved inside the Product Readiness evaluation environment. It is not evidence of real-world sales causality.
Re-test result recorded against the same test identity. Both states stay visible in history.
- Same shopper intent
- Same market
- Same relevant Brand Intent
- Comparable agent and test environment
- Versioned methodology
Where the environment or model materially changes, the result displays the new context rather than presenting the comparison as unchanged.
Compare like with like.
Results from different AI commerce environments may be informative, but they should not be read as identical experimental conditions. Every result carries the context needed to judge whether a comparison is fair.
| Environment | Version / profile | Date | Market | Test scope |
|---|---|---|---|---|
| Herm simulation environment | profile + version | per run | UK | 100 intents |
| Secondary commerce environment | profile + version | per run | UK | 100 intents |
| Market comparison run | profile + version | per run | DE | 100 intents |
Herm displays environment context rather than a single cross-environment ranking. False precision is worse than a missing number.
Tracked is not the same as certified.
ACP, UCP and other external standards are ecosystem specifications Herm tracks. Tracking a standard does not mean Herm certifies compliance with it or performs official validation on its behalf. Where Herm has a specific readiness test for a requirement, the result says so precisely.
Tracked as an ecosystem specification. Where Herm tests a specific requirement, the result names the requirement.
Tracked as an ecosystem specification. Tracking is not certification, compliance or official validation.
Product and retailer destinations are checked for validity in the tested market as part of Close.
What the measurement does not claim.
Eight statements Herm will not make about a Product Readiness result, listed here so nobody has to infer them.
A result is only interpretable against the methodology that produced it.
Historical results record the methodology and environment context needed to interpret them. No version numbers are published on this page until they are confirmed.
Same concepts, smaller scope.
The AI Commerce Readiness Test uses the same Product Readiness concepts, meaning shopper intents, Find / Answer / Trust / Close, evidence and Sellable Intent Rate, on a limited catalogue and test scope. It is a sample of the method, not the full subscription.
| Readiness Test | Product Readiness | |
|---|---|---|
| Shopper intents | limited panel | full panel, brand-informed |
| Catalogue scope | limited scope | full connected catalogue |
| Dimensions reported | Find · Answer · Trust · Close | Find · Answer · Trust · Close |
| Brand Intent | not configured | configured and evaluated |
| Root cause and re-testing | summary only | full diagnostics and re-tests |
Common questions
What exactly is being measured?
Whether an AI shopping agent can satisfy a representative customer buying situation using the product information the brand provides: retrieve an appropriate product, resolve the decisive questions, ground the recommendation in defensible evidence, and provide a valid path to purchase.
Does Herm read the model’s reasoning?
No. Herm evaluates observable shopping behaviour and evidence-backed commercial correctness. It does not read hidden model reasoning, and none is exposed in results.
Why are the weights not published?
The concepts, units and boundaries are public so a result can be interpreted and checked. The evaluator, the scoring formula, the weights, the thresholds and the intent-generation implementation are the product, and are not published.
Can we compare our score with another brand’s?
Not directly. A Sellable Intent Rate is only meaningful alongside its panel, market and environment. Two rates produced on different intents, markets or environments are not the same measurement.
Does a passing test guarantee an AI assistant will recommend us?
No. A passing Shopping Test does not guarantee an external AI service will always produce the same recommendation, and supported product information does not guarantee sales. The boundaries section lists the eight claims Herm will not make.
Is the Readiness Test the same measurement?
Same concepts, smaller scope. The AI Commerce Readiness Test runs a limited intent panel against a limited catalogue scope, without Brand Intent configured. It samples the method rather than replacing the full assessment.
Run the method on your own catalogue.
Connect a feed and see your first Shopping Tests, with the evidence behind every pass and fail.
Every definition on this page is the one used in the product.