Live Rung 02 · Be Sellable
Product Readiness · Be Sellable · Methodology

Know what “AI-ready to sell” actually means.

Product Readiness tests representative customer buying situations against your product catalogue and evaluates whether an AI shopping agent can retrieve an appropriate product, answer the customer’s questions, use trustworthy evidence and provide a valid path to purchase. This page explains what those measurements mean, how comparisons are made and what Herm deliberately does not claim.

Guiding principle: auditable methodology, protected implementation. Every concept, unit and boundary on this page is public. The evaluator, scoring and intent-generation implementations are not.

◍ herm · the measurement chain
  1. Shopper need A representative customer buying situation
  2. Agent behaviour Search, comparison, recommendation, explanation
  3. Observable outcome What the agent actually recommended and said
  4. Authoritative evidence + Brand Intent What the brand can support, and what it permits
  5. Herm evaluation Independent judgement of the outcome
  6. Find · Answer · Trust · Close Four dimensions, scored separately
  7. Pass / fail One tested intent, one result, evidence attached

Herm evaluates observable shopping behaviour and evidence-backed commercial correctness. It does not read hidden model reasoning.

01 The unit of measurement

Start with a buying situation, not a product field.

Product Readiness does not begin with your feed. It begins with a shopper intent: a representative customer buying situation describing a need, relevant constraints and the decision the shopper is trying to make. Everything that follows, meaning retrieval, evidence, correctness and destination, is judged against that situation.

◍ herm · shopper intent
Illustrative example

“I need a waterproof trail shoe under £150 for wide feet, suitable for long-distance walking in the UK.”

NeedTrail shoe
Use caseLong-distance walking
ConstraintsWaterproof · wide fit · under £150
MarketUK

A shopper intent is a measurement unit. One tested intent produces one Shopping Test result.

02 Where shopper intents come from

Outside-in testing, strengthened by first-party customer insight.

Herm can establish representative shopper situations from category and product context alone. Brands that contribute their own customer understanding make the panel closer to the customers they actually serve.

The better the test represents the customers the brand actually serves, the more commercially meaningful Product Readiness becomes.

Source A · Herm outside-in
Established from category and product context

Requires nothing from you beyond connected product information. This is what runs by default, and what makes a first result possible before any workshop.

Source B · Brand-provided understanding
Strengthened with what you already know
ICPsAudience researchCustomer interviewsBuying criteriaMerchandising insightCommon use casesObjections and questionsCategory research

Precision note. Tested intents are representative customer buying situations, not logs of literal user queries. Where an intent is derived from observed customer input, it is labelled as such in the result.

03 What the shopping agent receives

The agent is tested against the information the brand actually provides.

Product Readiness operates on connected commerce and product information. Herm normalizes connected information into a consistent test environment so that results are comparable between runs and between markets.

Supported input methods
  • Product-feed URL
  • XML
  • CSV
  • JSON
  • Google Merchant Center
  • API
Product information the test may use
ProductsVariantsAttributesPrice at test timeAvailability at test timeMarketsDestinationsApproved evidenceProduct relationships

Price and availability are evaluated as supplied by the connected source at test time. Herm does not independently verify live retailer inventory.

04 Agent and evaluator

The agent makes the recommendation. Herm judges it independently.

These are two separate systems doing two separate jobs. The shopping agent is not allowed to mark its own work.

Herm’s primary shopping simulation environment uses and extends Anthropic’s open-source Commerce Agents reference implementation as a credible commerce-agent architecture for product discovery, comparison and recommendation. Herm’s Product Readiness methodology, including shopper-intent evaluation, Brand Intent, evidence checking, root-cause analysis and re-testing, is Herm’s independent measurement layer.

Simulation environment · Shopping agent
searchescomparesrecommendsexplains

Behaves like a commerce agent a shopper could plausibly use. Its output is a recommendation and an explanation.

Independent measurement layer · Herm evaluator
checks suitabilitychecks evidencechecks Brand Intentchecks market validitychecks destination

Judges the recommendation against authoritative product evidence and applicable Brand Intent. Its output is a pass or a fail with the evidence attached.

Simulation environment attribution
not a partnershipnot a certificationnot an endorsementnot the live consumer experience of any assistant
05 The four dimensions

Find · Answer · Trust · Close

Every Shopping Test is evaluated on four dimensions. These are the canonical public definitions; the scoring behind each one is not published.

Dimension 01
Find
Retrieval and eligibility

Did the agent identify an appropriate and eligible product for the shopper’s need?

Observable considerations
Relevant categoryCorrect product or familyAppropriate variantMarket eligibilityShopper constraintsBetter eligible product missed

A technically retrievable product is not automatically the right product. A customer wants a wide-fit waterproof hiking shoe; the agent retrieves the standard-fit version. Find can fail even though the product family is broadly correct.

Dimension 02
Answer
Agent Answerability

Could the agent resolve the important parts of the customer’s buying decision using authoritative product information?

Observable considerations
Intended useSuitabilityCompatibilityProduct differenceFitConditionsConstraints

Field completeness and Answerability are not equivalent. A fully populated record can still leave the question that decides the purchase unresolved.

Dimension 03
Trust
Evidence and commercial validity

Was the recommendation grounded in product and commercial information the brand can stand behind?

Observable considerations
Product identityApproved claimsSupporting evidenceSupplied priceSupplied availabilityMarket informationProduct generationConsistency

Herm evaluates price and availability as supplied by the connected source at test time. Herm does not independently verify real-time retailer inventory.

Dimension 04
Close
Path to purchase

Can the shopper move from the recommendation to an appropriate and valid purchase destination?

Observable considerations
Valid product URLValid retailer URLMarket and geographySupplied availability

Close does not currently mean Herm completes checkout or processes a payment.

06 Why four dimensions instead of one score

The same failed journey can fail in very different ways.

Four tests. Four failures. Four completely different jobs for four different teams.

◍ herm · four failures
same journey, four ways to fail
How one failed journey resolves differently per dimension
DimensionWhat went wrongWhat it meansOwner
Find failure A better eligible product was missed. The agent recommended a product that works. It did not recommend the one that fits the constraints best. Product data / merchandising
Answer failure Correct product, unanswerable question. The right product was found, but the suitability question the shopper needed resolved could not be answered. Product information
Trust failure The recommendation used an unsupported claim. The agent justified its recommendation with something no authoritative product evidence establishes. Content / compliance
Close failure Correct recommendation, invalid destination. The recommendation was right and the purchase link was not valid for the UK market. Commerce / channels

A single overall score should never hide where the commercial failure occurred.

07 What gets reported

Two numbers, and neither one travels alone.

The Be Sellable Score is a summary Product Readiness measure derived from the underlying Find, Answer, Trust and Close evaluations. The underlying dimensions are always shown alongside it, because two brands with similar overall scores can have very different readiness problems.

The Sellable Intent Rate is the percentage of tested customer buying situations that result in an appropriate, evidence-backed and commercially valid product recommendation with a working path to purchase.

◍ herm · be sellable score
Definition · illustrative
72
Overall
Find 81
Answer 64
Trust 77
Close 66

Weights, formulas and thresholds are not published.

◍ herm · sellable intent rate
Primary success metric · illustrative calculation
Representative shopper intents tested 100
Sellable outcomes 84
Failed outcomes 16
84%
Sellable Intent Rate

Weights, formulas and thresholds are not published. The worked example below is an explanatory calculation, not customer proof. A rate without its panel is not a comparable figure.

08 Agent Answerability

How much of the buying decision can the evidence actually resolve?

Agent Answerability measures how much of the customer’s buying decision can be resolved using authoritative product information. Every important question in an intent resolves to one state.

Field completeness and Answerability are not equivalent. A record can be 100% populated and still leave the decisive question unanswered.

Supported

Authoritative evidence answers the question.

Unknown

The available authoritative information does not establish the answer.

Partial

Some but not all important aspects can be supported.

Conflict

Different authoritative or commercial information disagrees.

Not applicable

The question legitimately does not apply to this product.

◍ herm · evidence discipline
shopper questionSuitable for wide feet?
catalogueNo approved width or fit-suitability information.
herm reportsUnknown
herm never reportsNo

Evidence discipline: missing evidence is not a negative claim. Reporting “no” would be inventing a negative claim the brand never made. Unknown is a data gap you can close; “no” is a statement about the product.

09 Authoritative evidence

A recommendation needs evidence the brand can defend.

Product Readiness evaluates against brand-approved or otherwise authoritative product evidence. Every evaluation shows the evidence it relied on, so a result can be checked rather than trusted.

Public principle. The customer should be able to inspect what evidence supports the evaluation.

◍ herm · evidence the evaluation may draw on
Kinds of authoritative evidence and where they come from
EvidenceOrigin
Connected product data feed · API
Specifications brand supplied
Product documentation brand supplied
Approved claim libraries brand approved
Market-specific product information per market
Other customer-supplied authoritative evidence brand supplied

What Herm does not publish: how evidence is ranked, how conflicts between sources are resolved, and how source authority is weighted.

10 Brand Intent

Factual accuracy and commercial correctness are evaluated separately.

A recommendation can be entirely supported by product evidence and still be commercially wrong for the brand. Those are two different questions and Herm keeps them apart.

Product evidence asks: can authoritative product information establish the claim the agent made? Brand Intent asks: is this representation or recommendation one the brand wants made on its behalf?

◍ herm · worked example
Shopper intent Lightweight shoe for daily running
Product evidence The product technically functions for running
Brand Intent Do not position this product for running
Agent recommends it anyway Brand Intent failure

In the worked example nothing factually false was said. The Shopping Test still fails, and it fails as a Brand Intent failure rather than an evidence failure, because the fix belongs to merchandising rather than to product data.

11 What makes a test pass

Five conditions, all of them material.

A Shopping Test succeeds when the resulting recommendation satisfies the material requirements of the shopper intent, is supported by authoritative product information, respects applicable Brand Intent, is valid for the relevant market and provides an appropriate purchase destination.

Material requirements

The recommendation satisfies what the shopper intent actually required.

Authoritative support

The claims made are supported by product information the brand can defend.

Brand Intent

Applicable Brand Intent is respected.

Market validity

The recommendation is valid for the relevant market.

Purchase destination

An appropriate destination is provided.

A test fails when one or more material requirements are not satisfied. Materiality thresholds, confidence handling, evidence-sufficiency rules and candidate ranking are internal and are not published.

12 Failure classification

Eleven failure families.

Every failed Shopping Test is classified into one of these families, and each family maps to a different kind of fix. The classification is public; how failures are grouped and clustered is not.

◍ herm · failure families
11 families
The failure families a failed Shopping Test is classified into
FamilyDefinition
01 Retrieval No appropriate product was surfaced from the eligible catalogue.
02 Suitability A product was found but does not meet the shopper’s material constraints.
03 Answerability A decisive buying question could not be resolved from authoritative information.
04 Unsupported claim The recommendation rests on a statement no evidence establishes.
05 Brand Intent The recommendation conflicts with a brand rule about positioning or use.
06 Market eligibility The product is not valid or offered in the tested market.
07 Availability Supplied availability does not support the recommendation.
08 Destination No valid purchase destination, or a destination invalid for the market.
09 Variant The right family, the wrong variant for the stated constraints.
10 Product identity The product recommended cannot be reliably identified or matched.
11 Stale or inconsistent information Sources disagree, or the information no longer matches the product.
13 Root cause

A failed journey is evidence. A root cause makes it actionable.

Herm looks for recurring patterns across failed Shopping Tests to identify where a product, family or catalogue-wide information issue may be responsible. Root causes are described in customer-readable terms, not internal field names.

Signal
19 failed intents
Across three product families, all in the UK panel.
Pattern
Similar Answerability gap
The same kind of buying question goes unresolved each time.
Blast radius
214 products affected
Products sharing the same missing information.
Root cause
Intended-use information incomplete
Described in customer-readable terms, not internal field names.

Illustrative pattern. A root cause is a diagnosis of the tested evidence, not a claim about a single field in isolation.

14 Re-testing and comparability

A fix should survive the same test.

When relevant product information changes, Product Readiness can re-run the affected Shopping Tests. Re-testing shows whether performance improved inside the Product Readiness evaluation environment. It is not evidence of real-world sales causality.

◍ herm · same intent · run 12 · 14 Aug Before
same intent · same market · same brand intentWaterproof walking boots in a wide fit, no breaking in, under £180
result FAILED
reasonWide-fit evidence unavailable for the recommended variant.
agent selectedKessock Moorland Trek Mid GTX
price£139.00
Correct category
Budget respected
Wide-fit suitability
Break-in confirmation
Illustrative shopping test
◍ herm · same intent · run 13 · 22 Aug After
same intent · same market · same brand intentWaterproof walking boots in a wide fit, no breaking in, under £180
result PASSED
changeFit and waterproofing evidence added across 34 variants.
agent selectedKessock Moorland Trek Mid GTX Wide
price£139.00
Correct category
Budget respected
Wide-fit suitability
Break-in confirmation
Illustrative shopping test

Re-test result recorded against the same test identity. Both states stay visible in history.

Context preserved for a valid comparison
  • Same shopper intent
  • Same market
  • Same relevant Brand Intent
  • Comparable agent and test environment
  • Versioned methodology

Where the environment or model materially changes, the result displays the new context rather than presenting the comparison as unchanged.

15 Comparing environments

Compare like with like.

Results from different AI commerce environments may be informative, but they should not be read as identical experimental conditions. Every result carries the context needed to judge whether a comparison is fair.

◍ herm · environment context
carried with every result
Context recorded alongside each environment's results
EnvironmentVersion / profileDateMarketTest scope
Herm simulation environment profile + version per run UK 100 intents
Secondary commerce environment profile + version per run UK 100 intents
Market comparison run profile + version per run DE 100 intents

Herm displays environment context rather than a single cross-environment ranking. False precision is worse than a missing number.

16 Commerce standards

Tracked is not the same as certified.

ACP, UCP and other external standards are ecosystem specifications Herm tracks. Tracking a standard does not mean Herm certifies compliance with it or performs official validation on its behalf. Where Herm has a specific readiness test for a requirement, the result says so precisely.

Explore Commerce Channels →
ACP Tracked standard

Tracked as an ecosystem specification. Where Herm tests a specific requirement, the result names the requirement.

UCP Tracked standard

Tracked as an ecosystem specification. Tracking is not certification, compliance or official validation.

Feed and destination validity Tested

Product and retailer destinations are checked for validity in the tested market as part of Close.

17 Boundaries

What the measurement does not claim.

Eight statements Herm will not make about a Product Readiness result, listed here so nobody has to infer them.

A passing Shopping Test does not guarantee an external AI service will always produce the same recommendation.
A failed test does not prove a single product field caused the outcome unless the diagnostic supports it.
Product Readiness does not expose hidden model reasoning.
Supported product information does not guarantee sales.
Technical protocol conformity does not guarantee shopper success.
A correlation between a data change and an improved test result is not proof of external commercial causality.
Product Readiness does not currently execute payments or complete checkout.
Representative shopper intents are a measurement panel, not every possible customer request.
18 Methodology version

A result is only interpretable against the methodology that produced it.

Historical results record the methodology and environment context needed to interpret them. No version numbers are published on this page until they are confirmed.

◍ herm · methodology version
Version block · confirmation required
Current methodology version to confirm
Effective date date to confirm
Major changes summary to confirm
19 The AI Commerce Readiness Test

Same concepts, smaller scope.

The AI Commerce Readiness Test uses the same Product Readiness concepts, meaning shopper intents, Find / Answer / Trust / Close, evidence and Sellable Intent Rate, on a limited catalogue and test scope. It is a sample of the method, not the full subscription.

Learn about the AI Commerce Readiness Test →
◍ herm · scope comparison
AI Commerce Readiness Test compared with full Product Readiness
Readiness TestProduct Readiness
Shopper intents limited panel full panel, brand-informed
Catalogue scope limited scope full connected catalogue
Dimensions reported Find · Answer · Trust · Close Find · Answer · Trust · Close
Brand Intent not configured configured and evaluated
Root cause and re-testing summary only full diagnostics and re-tests

Common questions

What exactly is being measured?

Whether an AI shopping agent can satisfy a representative customer buying situation using the product information the brand provides: retrieve an appropriate product, resolve the decisive questions, ground the recommendation in defensible evidence, and provide a valid path to purchase.

Does Herm read the model’s reasoning?

No. Herm evaluates observable shopping behaviour and evidence-backed commercial correctness. It does not read hidden model reasoning, and none is exposed in results.

Why are the weights not published?

The concepts, units and boundaries are public so a result can be interpreted and checked. The evaluator, the scoring formula, the weights, the thresholds and the intent-generation implementation are the product, and are not published.

Can we compare our score with another brand’s?

Not directly. A Sellable Intent Rate is only meaningful alongside its panel, market and environment. Two rates produced on different intents, markets or environments are not the same measurement.

Does a passing test guarantee an AI assistant will recommend us?

No. A passing Shopping Test does not guarantee an external AI service will always produce the same recommendation, and supported product information does not guarantee sales. The boundaries section lists the eight claims Herm will not make.

Is the Readiness Test the same measurement?

Same concepts, smaller scope. The AI Commerce Readiness Test runs a limited intent panel against a limited catalogue scope, without Brand Intent configured. It samples the method rather than replacing the full assessment.

Run the method on your own catalogue.

Connect a feed and see your first Shopping Tests, with the evidence behind every pass and fail.

Every definition on this page is the one used in the product.