Index methodology · v1.0

Every number here can be checked.

The AI Visibility Index ranks brands by how five models answer real buying questions. This page publishes the panel, the question sets, the arithmetic and the limits — in enough detail that a disagreement can be settled with evidence rather than opinion.

Still to publish 9 outstanding

This page is not finished. These facts are not yet stated, and the page says so rather than filling the gap with something plausible. Each one blocks a production release until it is answered.

  • PanelPinned version string for each of the five models.
  • PanelThe date each model was pinned.
  • PanelRefresh day, time and timezone.
  • QuestionsWho drafts and who signs off a question set, and any outside reviewer.
  • ArithmeticRuns per question, per model, per week — and whether they are batched or spread.
  • ArithmeticConfirm the 60/40 split and the 4-position window match the scoring code.
  • ChangelogBackfill any changes made between launch and now.
  • CorrectionsCorrections email address.
  • CorrectionsFirst-response SLA, in business days.

What the score is not

The panel

Five models, equally weighted. Each is pinned to a specific version so a score is reproducible, and a re-pin is logged as a method change with the affected weeks flagged on the Index itself.

ModelVersion pinnedPinned onWeightAccess
ChatGPT [MODEL_VERSION] [PIN_DATE] 20% · equal API · no account history
Claude [MODEL_VERSION] [PIN_DATE] 20% · equal API · no account history
Gemini [MODEL_VERSION] [PIN_DATE] 20% · equal API · no account history
Perplexity [MODEL_VERSION] [PIN_DATE] 20% · equal API · no account history
Grok [MODEL_VERSION] [PIN_DATE] 20% · equal API · no account history
Refresh Weekly [REFRESH_DAY_TIME_TZ]

Language rule

Where a market runs in two languages, both runs are published. Models are stronger in English, and a local-language run can under-represent a strong local brand — showing only one would hide that.

TurkeyTurkish + English
United KingdomEnglish
NetherlandsDutch + English

The question sets

Questions come from how people actually phrase buying intent in a category. They are fixed before a sector goes live, and changing one is a method change.

01
Drafted from real shopping language

Questions come from how people phrase buying intent in that category — not from keyword volume, and never from a brand’s own terms.

02
Classified by intent and niche

Each question is tagged (recommendation, budget, gifting, comparison, authority, discovery) so a sector can be sliced without re-asking anything.

03
Frozen before launch

The set is fixed before the sector goes live. Adding or removing a question is a method change and goes in the changelog.

04
Re-reviewed quarterly

Language drifts. Sets are reviewed on a fixed cadence, never in response to a single brand’s result.

Who writes them [QUESTION_SET_OWNER]

The arithmetic

A model score combines how often a brand is named with how early it is named. A brand score is the equal-weight mean of the five.

Mention part 60 × (mentions ÷ questions asked)
Position part 40 × (1 − (average position − 1) ÷ 4)
Equal-weight mean Sum of the five model scores ÷ 5
Sampling [RUNS_PER_QUESTION]
Ties and absences A brand never named in a model’s answers scores zero for that model — it is not excluded from the mean. Excluding it would reward absence.

Worked example

An illustration with invented inputs, computed with the same formula the Index uses — the totals are derived, not typed.

ModelMentionsAvg posMentionPositionScore
ChatGPT 7 / 8 1.8 52.5 32.0 84.5
Claude 6 / 8 2.4 45.0 26.0 71.0
Gemini 6 / 8 2.2 45.0 28.0 73.0
Perplexity 8 / 8 1.4 60.0 36.0 96.0
Grok 6 / 8 2.3 45.0 27.0 72.0
Equal-weight mean of 5 396.5 ÷ 5 = 79.3

What we refuse to do

A ranking is only worth reading if the things that could corrupt it are named and ruled out.

No paid placement, ever

There is no advertising slot, no promoted row, and no premium that affects position. A customer and a non-customer are scored identically.

No advance notice of questions

Brands never see a sector’s question set before it runs. Neither do design partners.

No personalisation

Every run is from a clean context with no account, no history and no location signal.

No removing inconvenient models

All five models stay on the panel even when one disagrees sharply with the rest. Disagreement is data.

No retroactive smoothing

Published weeks are never quietly re-scored. A correction is a logged correction, with a date.

No score without transcripts

If the answer transcripts for a week aren’t stored, the week isn’t published.

Where this is weak

Stated plainly, because a methodology that only lists its strengths is marketing.

Thin sectors move more

A sector with few tracked questions or few listed brands is statistically noisier. Week-to-week deltas in small sectors should be read as direction, not magnitude.

Models are non-deterministic

The same question can return different brands on two runs. Averaging multiple runs reduces this but does not remove it.

A version pin is a snapshot

When a model is re-pinned, scores can shift for reasons that have nothing to do with your brand. Those weeks are flagged.

Mention is not endorsement

Being named early in an answer does not mean a model recommends you well — only that it recalls you. The text of the mention is not scored.

English bias is real

Models are stronger in English than in most other languages. A local-language run can under-represent a strong local brand, which is why both runs are published.

Changelog

Every model re-pin, formula change and correction, with a date. Weeks affected by a change are flagged on the Index itself, not only here.

Jul 27, 2026 launch Turkey opens with 8 live sectors on the five-model panel. Methodology v1.0 published.

Think a week is wrong?

Send it. Disputes are settled against the stored answer transcripts for that week, and a correction is published with a date rather than applied quietly.

Where to send it [CORRECTIONS_EMAIL]
First response within [CORRECTIONS_SLA]
Evidence used Stored answer transcripts for the week in question