How the AI Commerce Readiness Score is calculated
This methodology documents what Herm measures, how the four readiness rungs are evaluated, which inputs are required, what a score does and does not mean, and how the methodology changes over time.
What the score measures
The AI Commerce Readiness Score measures whether a brand is prepared for four different parts of AI-mediated commerce.
It is not one AI-visibility metric stretched across four labels. Each rung evaluates a different readiness question and uses the inputs appropriate to that question.
- Rung 01 · Be Visible Live
Does AI know and recommend the brand?
Brand Visibility measures the brand across Herm's fixed five-model panel, using the questions buyers in the relevant sector ask.
- Rung 02 · Be Sellable Live
Can AI use the brand's product data to complete a recommendation?
Product Readiness requires the brand to provide its product feed. Herm evaluates whether the data needed to quote and route the product is present, current and usable, by putting representative shopping situations through controlled Shopping Tests against the connected catalogue.
- Rung 03 · Be Relevant Early access
Can the brand's own signed-in experiences make a better choice for a connected shopper?
This rung is assessed from the Customer Intelligence connection and its eligibility rules. In an eligible signed-in session the brand sends the options it could already show, together with that shopper's connection context; Herm returns a ranking or a selection, and where supported an optional permitted explanation. The underlying shopper profile stays inside Herm. The rung therefore scores the brand's own in-session surfaces (product carousels, its own site assistant, offer slots) and does not cover email, CRM or any other offline channel, because decisions are returned only while a connected shopper is in session. It does not measure, and Herm does not influence, what a frontier AI assistant recommends to an individual shopper.
- Rung 04 · Be Preferred Live
Do the shoppers who buy the category buy the brand, and keep buying it?
This rung is measured from consented Herm receipt data: what share of verified category buyers also buy the brand, and whether they buy it again across retailers rather than in one store alone. Where receipt panel coverage is sufficient, currently Turkey, it requires no input from the brand, and an Offers connection adds redemption and repeat-purchase measurement. A verified receipt proves that a purchase occurred; cross-surface proof that one specific AI answer caused one specific purchase is not generally available today.
Scale on AI is the destination of the framework, not a fifth readiness sub-score. It does not contribute to the headline score.
The fixed five-model panel
- ChatGPT
- Claude
- Gemini
- Perplexity
- Grok
Brand Visibility measurement uses a fixed panel of ChatGPT, Claude, Gemini, Perplexity and Grok.
Keeping the panel fixed makes change over time interpretable. If the set of models being measured moved every cycle, a rise or fall in a brand’s visibility would be indistinguishable from a change in the measuring instrument.
Models are queried without a signed-in account and without personalisation, so a result reflects what the model returns to an unknown shopper rather than to a profiled one.
The questions the models are asked
The question set represents the kinds of commercial questions a shopper asks before choosing a product or provider: the questions that decide a shortlist, not questions designed to elicit a brand name.
Herm evaluates categories rather than prompting models to mention a specific brand. A prompt that names the brand tests recall of the prompt, not visibility in the category.
Question sets are built per sector and per country. Results measured under one sector’s question set are not comparable to another’s.
How Herm handles model variance
AI-model answers are probabilistic. The same question put to the same model twice can return different brands in a different order, and a single answer is not a stable ranking.
Herm therefore applies the same repeat-run and aggregation rules to every brand in a measured sector, rather than treating one response as the result. Brands within a sector are always measured in the same window, so a comparison is never between runs taken days apart.
A model that cannot complete a run is recorded as unavailable for that run rather than counted as an absence of the brand.
What happens when a rung is not connected
Some readiness inputs can be measured before a brand connects anything. Others require a brand-controlled connection, and Herm cannot assess them until it exists.
| Rung | Requirement | How it is assessed |
|---|---|---|
| Be Visible | No brand input required | Measured from Herm's five-model panel. |
| Be Sellable | Product feed required | The full Product Readiness assessment needs the brand’s product feed. |
| Be Relevant | Connection required | Assessed from the Customer Intelligence connection: an eligible signed-in brand-owned experience, and shoppers who have chosen to connect. Until both exist there is nothing to measure. |
| Be Preferred | No brand input required | Measured from Herm's consented receipt panel where receipt panel coverage is sufficient, currently Turkey. An Offers connection adds redemption and repeat-purchase measurement. Running Offers needs a brand account; being scored does not. |
An unavailable input is labelled unavailable. It is not presented as, and must not be read as, the same thing as a failing result: a brand with no connected product feed has not failed Be Sellable; it has not been measured on it.
Where Herm does not have enough measured dimensions to support a meaningful overall Readiness Score, the result is reported as a partial readiness assessment showing the dimensions that were measured, rather than as a headline score. An unavailable dimension is one Herm does not yet have the evidence to score; it is not a failed one.
What this version does not yet publish
This version documents what the score measures and how the inputs are gathered. It does not yet publish the per-rung input variables, their normalisation and weights, the headline-score formula, the number of prompts and repeat runs per model, or the refresh cadence for each component.
Those are production values, and publishing an approximation of them would be worse than publishing none: this page exists so a result can be checked, and a formula that does not match the one in use cannot be checked. They are documented in a future version, and the version history below records when.
Why Index and Brand Visibility results are not compared
The public sector Index and a brand's Brand Visibility result are separate measurements and are not interchangeable. The Index puts one shared question set to every brand in a sector, which is what makes its cross-brand ranking legitimate. A Brand Visibility result comes from questions designed around a single brand, so it is reported on its own terms and against that brand's own history. Herm does not compute a difference between the two, and no ranking on this site is derived from one applied to the other.
The public AI Visibility Index currently ranks 3,734 brands. The figure is read from the published Index when this page is built, so it moves as sectors are added.
What the Readiness Score does not mean
The Readiness Score is not a measure of brand quality, market share, revenue, customer satisfaction or the overall health of a business. It measures preparedness for AI-mediated commerce, and nothing else.
AI-model outputs are probabilistic and change as models, retrieval systems, sources and provider behaviour change. A score is a reading taken under a stated methodology at a stated time, not a permanent property of the brand.
A high Brand Visibility score does not mean AI can successfully transact with the brand. Product Readiness requires current product data and is assessed separately.
Brand Visibility is read from what a model returned: the answer text, the brands it named, the sources it cited and the explanation the answer itself gives. Herm does not observe a model’s internal reasoning, and no result on this site should be read as a statement about why a model produced an answer beyond what that answer states.
Product Readiness is read from controlled Shopping Tests run against the connected catalogue. Those tests are a controlled measurement environment. They are not the live consumer experience of ChatGPT, Claude or any other third-party assistant, they do not create carts or process payments, and a result on this rung is a statement about the brand’s readiness rather than evidence of what any external system did commercially.
A verified receipt proves that a purchase occurred. It does not, by itself, prove which AI answer caused the purchase. Cross-surface answer-level attribution requires an identifier that survives from the answer to the transaction, and that mechanism is not generally available today.
Version history
Herm may update this methodology as AI platforms, shopping behaviour and Herm’s own product capabilities evolve. Material changes are versioned so a score can be interpreted against the methodology in force when it was produced. Where a change materially breaks comparability with older scores, that is stated in the row.
| Version | Effective | Material change |
|---|---|---|
| 1.3 | Canonical definition and rung 03 scope corrected. The category definition previously named attribution as its fourth element; the fourth rung measures category preference and repeat purchase, and answer-level causal attribution is not available (see limitations). Rung 03 previously described scoring a brand’s CRM: Customer Intelligence returns decisions only inside an eligible signed-in session, so the rung covers the brand’s own in-session surfaces and no offline channel. Two measurement boundaries were added and one reporting rule was made explicit: a result with too few measured dimensions is reported as a partial readiness assessment rather than as a headline score. Measurement and scoring are unchanged. | |
| 1.2 | Sector benchmarking retired. The public Index and a Brand Visibility result are separate measurements on different question sets. Earlier versions stated they could be interpreted against one another; that was incorrect and no score was ever derived by comparing them. Measurement and scoring are unchanged. | |
| 1.1 | Rungs 03 and 04 renamed. 'Be Chosen' is now 'Be Relevant': the previous verb implied Herm influences frontier-model recommendations, which it does not. 'Be Proven' is now 'Be Preferred': the previous verb implied answer-level causal attribution, which is not available. Measurement and scoring are unchanged. | |
| 1.0 | Initial public Readiness methodology: scope of each rung, panel configuration, question-set construction, variance handling, connected/unconnected input rules, benchmarking and limitations. |
See the methodology applied to your brand.
Get the Readiness Score to see the four-rung framework applied to your own brand, with the underlying findings available in the product.