Watch AI shop your catalogue.
Give Herm a realistic customer need and watch an AI shopping agent search, compare and recommend from your catalogue. Herm evaluates whether it found the right product, supported the recommendation with authoritative information and sent the shopper to a valid place to buy.
A Shopping Test is one customer need, run end to end. Tests repeat on a schedule: read the methodology.
- 01 Shopper
- 02 Agent
- 03 Recommendation
- 04 Herm evaluation
“I’m going walking in Scotland for three days. I need something waterproof, comfortable for long distances, suitable for wide feet and below £150.”
Kessock Moorland Trek Mid GTX
“Waterproof with a GORE-TEX membrane and cushioned for long days on mixed terrain. Available in your size and in stock for UK delivery.”
Customers describe needs. Catalogues describe products. AI has to connect the two.
A Shopping Test starts with a customer, not a SKU. AI commerce requires the agent to translate a customer's situation into product criteria, retrieve eligible options and make a defensible choice. Product Readiness tests that chain.
Test real buying situations, not synthetic field checks.
Herm builds representative buying situations from the outside in: the way a shopper describes a need, with the constraints and questions they bring with them.
Testing gets stronger when the brand contributes what it already knows about its customers. Intents are reviewable and editable before a run.
Category and buying-situation intelligence built independently of the brand.
What the brand already knows about who buys and why.
These are representative buying situations. Generation method and underlying datasets are proprietary.
Customers ask harder questions than “find me a red shoe”.
A suite mixes decision types, because the ways a catalogue fails a customer are not all the same. These are the types currently supported.
| Decision type | Representative buying situation | What the test establishes |
|---|---|---|
| Need matching | “What is the best option for long-distance walking?” | Whether an appropriate eligible product is retrieved and defended for the stated need. |
| Constraint handling | “Under £150, wide feet, delivered in the UK.” | Whether every stated constraint is respected, including price, fit and market. |
| Product comparison | “Which of these two is better for ankle stability?” | Whether the comparison is supported by authoritative product information. |
| Compatibility | “Will these work with the gaiters I already own?” | Whether compatibility can be established from the product record. |
| Trade-off | “I would rather have comfort than low weight.” | Whether the agent resolves competing priorities in the customer's favour. |
| Current-model choice | “Is this the latest version of the boot?” | Whether the current generation is recommended over an obsolete predecessor. |
| Market eligibility | “I am ordering from the UK.” | Whether the recommended product and destination are appropriate for the shopper's market. |
What every Shopping Test asks, in order.
Find, Answer, Trust and Close are the four things that have to hold for one customer need to end well. A test that clears three of them still ends with a shopper who cannot buy the right thing.
Did the agent retrieve an appropriate eligible product?
- wrong category
- wrong variant
- a better product missed
- market-ineligible item retrieved
Could it answer the customer's important questions?
- product information does not establish whether this model suits wide feet
Agent Answerability lives here.
Was the recommendation grounded in approved and current information?
- unsupported product claim
- ambiguous product identity
- incorrect supplied price
- availability mismatch
- unapproved positioning
Could the shopper reach a valid purchase destination?
- invalid product URL
- invalid retailer URL
- supplied availability wrong
- wrong geography
Close covers reaching a valid place to buy: a working product URL, a retailer destination, supplied availability and the correct geography. Herm does not create carts, execute checkout or process payment.
See what the agent considered, and why the final choice passed or failed.
Reading only the final answer hides the more useful failure: the product that should have won and was never retrieved. Select a candidate to see the observed action and Herm's independent view of it.
Retrieved in the first search, then set aside as above the stated price ceiling.
Rejection correct. £169.00 exceeds the £150 constraint the customer stated.
A confident answer can still be commercially wrong.
Agents write well. Fluency is not the success metric. Herm classifies what was wrong about a recommendation, which is what tells a product team where to work.
The agent states the boot is waterproof without authoritative support in the product record.
TrustA technical climbing product recommended for a three-day walking trip.
Find · Brand IntentA product recommended that is not available in the shopper's geography.
Find · CloseCorrect boot family, standard fitting recommended to a customer who asked for wide.
FindThe product choice is right and the purchase destination is broken or out of market.
CloseThe catalogue is not the only source of truth.
Brands may have commercial, regulatory or positioning rules that determine whether a recommendation is acceptable even when the product attributes alone do not capture them.
Brand Intent lets you state what correct means for your brand. Every Shopping Test is then evaluated against it as well as the product record.
Every failed test shows what could be supported, and what could not.
A verdict you cannot inspect is not worth acting on. Each claim in the recommendation resolves to a product field and the source it came from. Select a row to follow the chain.
The agent asserted waterproofing and the assertion holds against authoritative product information.
Assertions are compared with authoritative product information supplied by the brand. Evidence selection is proprietary.
One bad recommendation may reveal a catalogue-wide problem.
Herm groups recurring failure patterns so teams can see whether a problem belongs to one product, one family or the whole catalogue.
State intended activity and fit width authoritatively on the affected product records, with evidence.
Fix the product information. Ask the same customer again.
Same shopper need, same category and market, same evaluation. Re-run once the product record changes.
Product data updated: 214 products, fit width and intended activity approved 22 Aug, re-tested automatically.
This is measured improvement inside the Product Readiness evaluation. It is not a claim about external sales.
The catalogue changes. The test suite keeps asking.
Product Readiness is continuous assurance, not a launch certification. A test that passed last month stops passing when a product, a price, a destination or the agent environment moves.
Regression. Something changed and a customer need stopped being served.
Recovery. The fix held when the same shopper came back.
- 24 Jul New product introduced Kessock Moorland Trek Mid GTX added to the suite baseline
- 28 Jul Availability changed Cairn Summit out of stock · UK passed to failed
- 02 Aug Destination changed retailer URL redirect broke test 07 passed to failed
- 09 Aug Product information changed 34 variants updated in the connected source
- 10 Aug Regression detected 3 tests · walking & hiking footwear passed to failed
- 14 Aug Suite re-run 50 tests · 19 failed
- 22 Aug Product enriched, tests re-run fit width + intended activity approved · 214 products failed to passed
- 30 Aug Protocol environment changed feed requirement update · re-run, result restored restored
Built on a real commerce-agent reference architecture.
Herm's primary simulation environment uses and extends Anthropic's open-source Commerce Agents reference implementation. Herm adds its own catalogue representation, shopper-intent testing, Brand Intent, independent evaluation, diagnosis and re-testing.
Shopping Tests are framework-independent. The same shopper need can be run against more than one agent environment.
Anthropic is not affiliated with or endorsing Herm. A test environment is not the live consumer experience of any assistant.
A recommendation often fails because the product information cannot resolve the customer's question.
Agent Answerability maps which decision-relevant questions your catalogue can support, family by family, before a shopper ever asks one.
Common questions
What exactly is one Shopping Test?
One representative customer buying situation, run end to end: the need, the agent searching and comparing your catalogue, the recommendation it makes, and an independent evaluation of whether that recommendation served the customer.
Does Herm see the model’s reasoning?
No. Herm evaluates observable behaviour: the tool actions the agent took, the products it retrieved and rejected, and the recommendation and explanation it produced. Hidden model reasoning is not read and is not shown.
Who decides whether a test passed?
Herm, independently of the agent. The shopping agent produces the recommendation; a separate evaluation layer judges it against authoritative product evidence and any applicable Brand Intent. The agent is never allowed to mark its own work.
Are these real customer queries?
They are representative buying situations, not logs of literal user queries. Where an intent is derived from observed customer input, the result labels it as such.
Can we run the same test against different AI systems?
Yes. Shopping Tests are framework-independent, so the same shopper need can be run against more than one agent environment. Results carry their environment context, because different environments are not identical experimental conditions.
See what happens when AI shops your catalogue.
One realistic customer need, run end to end, with the evidence behind the verdict.
Sample figures throughout.