Index methodology · v1.0

Every number here can be checked.

The AI Visibility Index ranks brands by how five models answer real buying questions. This page publishes the panel, the question sets, the arithmetic and the limits, in enough detail that a disagreement can be settled with evidence rather than opinion.

What the score is not

The panel

Five models, equally weighted. Each is pinned to a specific version so a score is reproducible, and a re-pin is logged as a method change with the affected weeks flagged on the Index itself.

ModelVersion in useFirst measuredWeightAccess
ChatGPT openai/gpt-5.6-luna-pro May 25, 2026 20% · equal API · no account history
Claude anthropic/claude-haiku-4.5 May 25, 2026 20% · equal API · no account history
Gemini google/gemini-3.6-flash May 25, 2026 20% · equal API · no account history
Perplexity perplexity/sonar Jun 12, 2026 20% · equal API · no account history
Grok x-ai/grok-4.3 May 25, 2026 20% · equal API · no account history
Refresh One sector per market per week · every sector once per its market's cycle Mondays at 10:00 GMT

Language rule

Where a market runs in two languages, both runs are published. Models are stronger in English, and a local-language run can under-represent a strong local brand. Showing only one would hide that.

TurkeyTurkish + English
United KingdomEnglish
NetherlandsDutch + English

The question sets

Questions come from how people actually phrase buying intent in a category. They are fixed before a sector goes live, and changing one is a method change.

01
Drafted from real shopping language

Questions come from how people phrase buying intent in that category, not from keyword volume, and never from a brand’s own terms.

02
Classified by intent and niche

Each question is tagged (recommendation, budget, gifting, comparison, authority, discovery) so a sector can be sliced without re-asking anything.

03
Frozen before launch

The set is fixed before the sector goes live. Adding or removing a question is a method change and goes in the changelog.

04
Re-reviewed quarterly

Language drifts. Sets are reviewed on a fixed cadence, never in response to a single brand’s result.

Who writes them Drafted and signed off in-house by the Herm research team. No outside reviewer.

The arithmetic

A score combines three things: how often a brand is named, how early it is named, and how many of the models name it at all. Each part is measured against the strongest ranked brand in the same cut, so 100 is the leader rather than a theoretical ceiling.

Mention part 45% × (mention rate ÷ the highest mention rate among ranked brands)
Position part 30% × (MRR ÷ the highest MRR among ranked brands). MRR is the mean of 1 ÷ the rank of a brand’s first appearance, over the answers that name it.
Breadth part 25% × (models that named the brand ÷ models in the selection)
What 100 means Each part is scaled against the strongest ranked brand in the cut, so the leader reads 100. Entities the sector does not rank are still measured and shown separately, and because the scale is not set by them, one can read above 100.
Per-model columns A single-model column drops the breadth part, which is always the same for one model, and reweights the other two to 60/40. A column is its own calculation, not a share of the headline, so the five do not average to it.
Sampling 5 runs per question, per model, collected in one window
Ties and absences A brand never named in a model’s answers has no mention rate and no MRR for that model, so it scores zero there. Breadth already counts the models that did name it.

Worked example

The current leader of a live sector, computed by the same engine that builds the rankings. Every figure is derived, not typed.

PartInputScaled
Mention · 45% 52.89% of answers 100.0
Position · 30% MRR 0.5464 81.0
Breadth · 25% 5 of 5 models 100.0
Gazelle · bicycles-ebikes · September 2026 94.3

What we refuse to do

A ranking is only worth reading if the things that could corrupt it are named and ruled out.

No paid placement, ever

There is no advertising slot, no promoted row, and no premium that affects position. A customer and a non-customer are scored identically.

No advance notice of questions

Brands never see a sector’s question set before it runs. Neither do design partners.

No personalisation

Every run is from a clean context with no account, no history and no location signal.

No removing inconvenient models

All five models stay on the panel even when one disagrees sharply with the rest. Disagreement is data.

No retroactive smoothing

Published weeks are never quietly re-scored. A correction is a logged correction, with a date.

No score without transcripts

If the answer transcripts for a week aren’t stored, the week isn’t published.

Where this is weak

Stated plainly, because a methodology that only lists its strengths is marketing.

Thin sectors move more

A sector with few tracked questions or few listed brands is statistically noisier. Week-to-week deltas in small sectors should be read as direction, not magnitude.

Models are non-deterministic

The same question can return different brands on two runs. Averaging multiple runs reduces this but does not remove it.

A version pin is a snapshot

When a model is re-pinned, scores can shift for reasons that have nothing to do with your brand. Those weeks are flagged.

Mention is not endorsement

Being named early in an answer does not mean a model recommends you well, only that it recalls you. The text of the mention is not scored.

English bias is real

Models are stronger in English than in most other languages. A local-language run can under-represent a strong local brand, which is why both runs are published.

Changelog

Every model re-pin, formula change and correction, with a date. Weeks affected by a change are flagged on the Index itself, not only here.

Jul 27, 2026 launch Turkey opens on the five-model panel. Methodology v1.0 published.

Think a week is wrong?

Send it. Disputes are settled against the stored answer transcripts for that week, and a correction is published with a date rather than applied quietly.

Where to send it team@herm.io
First response within 5 business days
Evidence used Stored answer transcripts for the week in question