Every number here can be checked.
The AI Visibility Index ranks brands by how five models answer real buying questions. This page publishes the panel, the question sets, the arithmetic and the limits — in enough detail that a disagreement can be settled with evidence rather than opinion.
This page is not finished. These facts are not yet stated, and the page says so rather than filling the gap with something plausible. Each one blocks a production release until it is answered.
- PanelPinned version string for each of the five models.
- PanelThe date each model was pinned.
- PanelRefresh day, time and timezone.
- QuestionsWho drafts and who signs off a question set, and any outside reviewer.
- ArithmeticRuns per question, per model, per week — and whether they are batched or spread.
- ArithmeticConfirm the 60/40 split and the 4-position window match the scoring code.
- ChangelogBackfill any changes made between launch and now.
- CorrectionsCorrections email address.
- CorrectionsFirst-response SLA, in business days.
What the score is not
- Not a traffic or revenue estimate — it says nothing about clicks or sales.
- Not a quality judgement of your product. Models can be confidently wrong.
- Not personalised. No shopper history, no location, no account.
- Not purchasable. There is no paid tier of the Index, in any market.
The panel
Five models, equally weighted. Each is pinned to a specific version so a score is reproducible, and a re-pin is logged as a method change with the affected weeks flagged on the Index itself.
Language rule
Where a market runs in two languages, both runs are published. Models are stronger in English, and a local-language run can under-represent a strong local brand — showing only one would hide that.
The question sets
Questions come from how people actually phrase buying intent in a category. They are fixed before a sector goes live, and changing one is a method change.
Questions come from how people phrase buying intent in that category — not from keyword volume, and never from a brand’s own terms.
Each question is tagged (recommendation, budget, gifting, comparison, authority, discovery) so a sector can be sliced without re-asking anything.
The set is fixed before the sector goes live. Adding or removing a question is a method change and goes in the changelog.
Language drifts. Sets are reviewed on a fixed cadence, never in response to a single brand’s result.
The arithmetic
A model score combines how often a brand is named with how early it is named. A brand score is the equal-weight mean of the five.
Worked example
An illustration with invented inputs, computed with the same formula the Index uses — the totals are derived, not typed.
What we refuse to do
A ranking is only worth reading if the things that could corrupt it are named and ruled out.
There is no advertising slot, no promoted row, and no premium that affects position. A customer and a non-customer are scored identically.
Brands never see a sector’s question set before it runs. Neither do design partners.
Every run is from a clean context with no account, no history and no location signal.
All five models stay on the panel even when one disagrees sharply with the rest. Disagreement is data.
Published weeks are never quietly re-scored. A correction is a logged correction, with a date.
If the answer transcripts for a week aren’t stored, the week isn’t published.
Where this is weak
Stated plainly, because a methodology that only lists its strengths is marketing.
A sector with few tracked questions or few listed brands is statistically noisier. Week-to-week deltas in small sectors should be read as direction, not magnitude.
The same question can return different brands on two runs. Averaging multiple runs reduces this but does not remove it.
When a model is re-pinned, scores can shift for reasons that have nothing to do with your brand. Those weeks are flagged.
Being named early in an answer does not mean a model recommends you well — only that it recalls you. The text of the mention is not scored.
Models are stronger in English than in most other languages. A local-language run can under-represent a strong local brand, which is why both runs are published.
Changelog
Every model re-pin, formula change and correction, with a date. Weeks affected by a change are flagged on the Index itself, not only here.
Think a week is wrong?
Send it. Disputes are settled against the stored answer transcripts for that week, and a correction is published with a date rather than applied quietly.