Live Rung 02 · Be Sellable
Product Readiness · Be Sellable

Watch AI shop your catalogue.

Give Herm a realistic customer need and watch an AI shopping agent search, compare and recommend from your catalogue. Herm evaluates whether it found the right product, supported the recommendation with authoritative information and sent the shopper to a valid place to buy.

A Shopping Test is one customer need, run end to end. Tests repeat on a schedule: read the methodology.

◍ herm · shopping test 31 LIVE
walking & hiking footwear · UK · herm test profile · claude
  1. 01 Shopper
  2. 02 Agent
  3. 03 Recommendation
  4. 04 Herm evaluation
Shopper need

“I’m going walking in Scotland for three days. I need something waterproof, comfortable for long distances, suitable for wide feet and below £150.”

waterproofmulti-day walkinglong distancewide fitunder £150UK
Agent · observed actions
26 products retrieved · 4 shortlisted
Requirements read from the customer need
waterproof · multi-day walking · wide fit · ≤ £150 · UK
Interpreted
Catalogue queried for walking footwear
26 products retrieved · 4 relevant to this need
Searched
Kessock Cairn Summit Mid GTX
£169.00 · rejected: above stated budget
Considered
Kessock Trailhead Low
US-only destination · rejected: market not eligible
Rejected
Kessock Moorland Trek Mid GTX vs Kessock Cairn Summit Mid GTX
waterproof membrane · weight · midsole · price
Compared
Kessock Moorland Trek Mid GTX
£139.00 · in stock · UK destination supplied
Recommended
Observable tool actions and responses only. Herm does not read hidden model reasoning.
Recommendation

Kessock Moorland Trek Mid GTX

£139.00

“Waterproof with a GORE-TEX membrane and cushioned for long days on mixed terrain. Available in your size and in stock for UK delivery.”

variantsize 9 · standard fit
availabilityin stock · UK
kessock.co.uk/p/moorland-trek-mid-gtx
Herm evaluation
independent of the agent
Category fit
Budget
Market eligibility
Waterproof evidence
Wide-fit evidence
Long-distance suitability
Availability
Valid destination
Result FAILED
Root issue
Wide-fit suitability cannot be established from authoritative product information.
stage 4 of 4 · herm evaluation
Live product · sample workspace
01 Two languages

Customers describe needs. Catalogues describe products. AI has to connect the two.

A Shopping Test starts with a customer, not a SKU. AI commerce requires the agent to translate a customer's situation into product criteria, retrieve eligible options and make a defensible choice. Product Readiness tests that chain.

◍ herm · two languages
Catalogue language
membranewaterproof
weight342 g
midsoleEVA
fitstandard
price£139
Structured. Accurate. Written for systems.
Shopper language
rainy hiking tripwalking all daywide feetbeginnerunder £150
Situational. Constrained. Never uses your field names.
02 Shopper intents

Test real buying situations, not synthetic field checks.

Herm builds representative buying situations from the outside in: the way a shopper describes a need, with the constraints and questions they bring with them.

Testing gets stronger when the brand contributes what it already knows about its customers. Intents are reviewable and editable before a run.

◍ herm · shopper intents
Herm outside-in

Category and buying-situation intelligence built independently of the brand.

Brand first-party

What the brand already knows about who buys and why.

Representative shopper intents
31 Three days walking in Scotland. Waterproof, comfortable over long distances, wide feet, under £150.
07 First proper walking boots for weekend hikes in the Peak District. Not sure about sizing.
44 Something lighter than my current boots for fast day walks, but still waterproof.
50 intents in this suite · reviewed before run 12
Strengthen intents with
ICP definitionsAudience researchCustomer insightsCategory researchCommon buying questionsUse casesMerchandising knowledge

These are representative buying situations. Generation method and underlying datasets are proprietary.

03 Test types

Customers ask harder questions than “find me a red shoe”.

A suite mixes decision types, because the ways a catalogue fails a customer are not all the same. These are the types currently supported.

◍ herm · decision types
7 supported types
Supported Shopping Test decision types
Decision typeRepresentative buying situationWhat the test establishes
Need matching “What is the best option for long-distance walking?” Whether an appropriate eligible product is retrieved and defended for the stated need.
Constraint handling “Under £150, wide feet, delivered in the UK.” Whether every stated constraint is respected, including price, fit and market.
Product comparison “Which of these two is better for ankle stability?” Whether the comparison is supported by authoritative product information.
Compatibility “Will these work with the gaiters I already own?” Whether compatibility can be established from the product record.
Trade-off “I would rather have comfort than low weight.” Whether the agent resolves competing priorities in the customer's favour.
Current-model choice “Is this the latest version of the boot?” Whether the current generation is recommended over an obsolete predecessor.
Market eligibility “I am ordering from the UK.” Whether the recommended product and destination are appropriate for the shopper's market.
04 The evaluation

What every Shopping Test asks, in order.

Find, Answer, Trust and Close are the four things that have to hold for one customer need to end well. A test that clears three of them still ends with a shopper who cannot buy the right thing.

Find 01

Did the agent retrieve an appropriate eligible product?

Fails when
  • wrong category
  • wrong variant
  • a better product missed
  • market-ineligible item retrieved
Answer 02

Could it answer the customer's important questions?

Fails when
  • product information does not establish whether this model suits wide feet

Agent Answerability lives here.

Trust 03

Was the recommendation grounded in approved and current information?

Fails when
  • unsupported product claim
  • ambiguous product identity
  • incorrect supplied price
  • availability mismatch
  • unapproved positioning
Close 04

Could the shopper reach a valid purchase destination?

Fails when
  • invalid product URL
  • invalid retailer URL
  • supplied availability wrong
  • wrong geography

Close covers reaching a valid place to buy: a working product URL, a retailer destination, supplied availability and the correct geography. Herm does not create carts, execute checkout or process payment.

05 Consideration

See what the agent considered, and why the final choice passed or failed.

Reading only the final answer hides the more useful failure: the product that should have won and was never retrieved. Select a candidate to see the observed action and Herm's independent view of it.

◍ herm · candidates
26 retrieved · 4 relevant to this need
Herm finding: more appropriate eligible product not considered. The wide fitting exists in the catalogue and is not expressed in a form the agent can retrieve or defend.
Candidate detail
Kessock Cairn Summit Mid GTX Rejected
£169.00 · standard fit · in stock · UK
Observed agent action

Retrieved in the first search, then set aside as above the stated price ceiling.

Herm evaluation

Rejection correct. £169.00 exceeds the £150 constraint the customer stated.

Product record
price 169.00 GBP
fit standard
availability in_stock · GB
Live product · sample workspace
06 Correctness

A confident answer can still be commercially wrong.

Agents write well. Fluency is not the success metric. Herm classifies what was wrong about a recommendation, which is what tells a product team where to work.

Factually wrong
Unsupported claim

The agent states the boot is waterproof without authoritative support in the product record.

Trust
Commercially wrong
Wrong use case

A technical climbing product recommended for a three-day walking trip.

Find · Brand Intent
Eligibility wrong
Wrong market

A product recommended that is not available in the shopper's geography.

Find · Close
Variant wrong
Right family, wrong item

Correct boot family, standard fitting recommended to a customer who asked for wide.

Find
Destination wrong
Nowhere to buy

The product choice is right and the purchase destination is broken or out of market.

Close
07 Brand Intent

The catalogue is not the only source of truth.

Brands may have commercial, regulatory or positioning rules that determine whether a recommendation is acceptable even when the product attributes alone do not capture them.

Brand Intent lets you state what correct means for your brand. Every Shopping Test is then evaluated against it as well as the product record.

Explore Brand Intent →
◍ herm · brand intent
4 active · applied to every test
Never recommend Cairn Summit for road running. Product rule · 1 family On
Sustainability claims require approved evidence. Claim rule · catalogue-wide On
Prefer the latest generation where alternatives otherwise fit equally. Merchandising rule · catalogue-wide On
UK customers must not receive US-only destinations. Market rule · UK / EN On
Rules are authored by the brand. Herm reports where an agent's behaviour conflicts with them.
08 Evidence

Every failed test shows what could be supported, and what could not.

A verdict you cannot inspect is not worth acting on. Each claim in the recommendation resolves to a product field and the source it came from. Select a row to follow the chain.

◍ herm · evidence · Kessock Moorland Trek Mid GTX
5 assertions checked
Evidence chain · 3 levels
Waterproof Supported
GORE-TEX · 3-layer
Evidence chain
Evaluation trust · claim supported
Field waterproof_membrane = gore-tex, 3-layer
Source approved product specification · v7 · 12 Jun 2026
What it means

The agent asserted waterproofing and the assertion holds against authoritative product information.

Live product · sample workspace

Assertions are compared with authoritative product information supplied by the brand. Evidence selection is proprietary.

09 Diagnosis

One bad recommendation may reveal a catalogue-wide problem.

Herm groups recurring failure patterns so teams can see whether a problem belongs to one product, one family or the whole catalogue.

◍ herm · root cause
run 12 · 3 groups
Root issue
intended_activity
incomplete or ambiguous
Affected
214
products
Observed impact
6
product families
Recommended action

State intended activity and fit width authoritatively on the affected product records, with evidence.

fit width not published · 96 productsdestination geography mismatch · 12 products
Live product · sample workspace
10 Re-test

Fix the product information. Ask the same customer again.

Same shopper need, same category and market, same evaluation. Re-run once the product record changes.

◍ herm · product record changes
214 products · approved 22 Aug
intended_activity Multi-day hill and mountain walking on mixed terrain. Cushioned for consecutive long days. approved product specification · 214 products
fit_width Standard and wide fittings published per variant, with last width in mm. product page size table · promoted to variant record
variant_record Wide variant exposed with its own identifier and destination. connected catalogue · re-published 22 Aug
◍ herm · test 31 · run 12 · 14 Aug Before
shopper need · unchangedWaterproof walking footwear for three days in Scotland, long distances, wide feet, under £150
result FAILED
reasonWide fitting exists on the product page and not in the variant record.
agent selectedKessock Moorland Trek Mid GTX
price£139.00 · standard fit
Category fit
Budget
Waterproof evidence
Wide-fit evidence
Long-distance suitability
Valid destination
Live product · sample workspace
◍ herm · test 31 · run 13 · 22 Aug After
shopper need · unchangedWaterproof walking footwear for three days in Scotland, long distances, wide feet, under £150
result PASSED
changeAppropriate product selected with an evidence-backed fit explanation.
agent selectedKessock Moorland Trek Mid GTX Wide
price£139.00 · wide fit
Category fit
Budget
Waterproof evidence
Wide-fit evidence
Long-distance suitability
Valid destination
Live product · sample workspace

Product data updated: 214 products, fit width and intended activity approved 22 Aug, re-tested automatically.

This is measured improvement inside the Product Readiness evaluation. It is not a claim about external sales.

11 Continuous

The catalogue changes. The test suite keeps asking.

Product Readiness is continuous assurance, not a launch certification. A test that passed last month stops passing when a product, a price, a destination or the agent environment moves.

new productsproduct information changedavailability changeddestination changedprotocol environment changedregressions
Passed to Failed

Regression. Something changed and a customer need stopped being served.

Failed to Passed

Recovery. The fix held when the same shopper came back.

◍ herm · test history
walking & hiking footwear · 5 weeks
  1. 24 Jul New product introduced Kessock Moorland Trek Mid GTX added to the suite baseline
  2. 28 Jul Availability changed Cairn Summit out of stock · UK passed to failed
  3. 02 Aug Destination changed retailer URL redirect broke test 07 passed to failed
  4. 09 Aug Product information changed 34 variants updated in the connected source
  5. 10 Aug Regression detected 3 tests · walking & hiking footwear passed to failed
  6. 14 Aug Suite re-run 50 tests · 19 failed
  7. 22 Aug Product enriched, tests re-run fit width + intended activity approved · 214 products failed to passed
  8. 30 Aug Protocol environment changed feed requirement update · re-run, result restored restored
Live product · sample workspace
12 Where tests run

Built on a real commerce-agent reference architecture.

Herm's primary simulation environment uses and extends Anthropic's open-source Commerce Agents reference implementation. Herm adds its own catalogue representation, shopper-intent testing, Brand Intent, independent evaluation, diagnosis and re-testing.

Shopping Tests are framework-independent. The same shopper need can be run against more than one agent environment.

◍ herm · environments
Live environments
Herm test profile · Claude extends the open-source Commerce Agents reference LIVE
Herm test profile · OpenAI shopping and product discovery behaviour LIVE
Herm test profile · Gemini shopping and product discovery behaviour LIVE
Roadmap
Shopify-specific agent environments not available today ○ Roadmap
Custom merchant agents not available today ○ Roadmap

Anthropic is not affiliated with or endorsing Herm. A test environment is not the live consumer experience of any assistant.

13 Next

A recommendation often fails because the product information cannot resolve the customer's question.

Agent Answerability maps which decision-relevant questions your catalogue can support, family by family, before a shopper ever asks one.

Explore Agent Answerability →

Common questions

What exactly is one Shopping Test?

One representative customer buying situation, run end to end: the need, the agent searching and comparing your catalogue, the recommendation it makes, and an independent evaluation of whether that recommendation served the customer.

Does Herm see the model’s reasoning?

No. Herm evaluates observable behaviour: the tool actions the agent took, the products it retrieved and rejected, and the recommendation and explanation it produced. Hidden model reasoning is not read and is not shown.

Who decides whether a test passed?

Herm, independently of the agent. The shopping agent produces the recommendation; a separate evaluation layer judges it against authoritative product evidence and any applicable Brand Intent. The agent is never allowed to mark its own work.

Are these real customer queries?

They are representative buying situations, not logs of literal user queries. Where an intent is derived from observed customer input, the result labels it as such.

Can we run the same test against different AI systems?

Yes. Shopping Tests are framework-independent, so the same shopper need can be run against more than one agent environment. Results carry their environment context, because different environments are not identical experimental conditions.

See what happens when AI shops your catalogue.

One realistic customer need, run end to end, with the evidence behind the verdict.

Sample figures throughout.