AI Index / Methodology

How the Index is measured

The AI Recommendation Index measures what AI assistants actually recommend when buyers ask. Every number on an Index page traces back to the process below — and every page discloses its own sample size and window. The measurement parameters on this page (models, gates, guardrails) are read live from the pipeline that collects the data, not hand-maintained copy.

§ Prompt design

Each category carries a fixed set of buyer-shaped prompts — 6–8 by design; the Statistical robustness section reads the live distribution: the questions real buyers type, like “best [category] for [audience]”, “[category] alternatives”, or “what should I use for [job]”.

Prompts are strictly neutral — they never contain product names, never steer toward or away from any vendor, and stay stable between samples so trends compare like with like.

§ Engines & measured models

Categories are measured across up to 5 engines via their official APIs, with web search enabled where the engine supports it. An AI engine is not one thing — answers depend on the exact model behind it, so the Index discloses the model measured per engine:

  • CHATGPT (OpenAI) — model: chat-latest · changeover in progress — published numbers count only chat-latest runs
  • CLAUDE (Anthropic) — model: claude-sonnet-5 · changeover in progress — published numbers count only claude-sonnet-5 runs
  • GEMINI (Google) — model: gemini-2.5-flash
  • PERPLEXITY (Perplexity) — model: sonar
  • GROK (xAI) — model: grok-4.20-0309-non-reasoning

Each category page shows per-engine shares separately — engines disagree, and that disagreement is part of the data. Not every category runs on every engine; the engines measured are listed per category.

§ Sampling, windows & the minimum-sample gate

Sampling runs continuously, spread across the batch cycle to avoid time-of-day artifacts. At the current measurement budget, each engine samples each category roughly monthly — cadence scales with budget, and every published number discloses exactly how many sampled answers it rests on.

Published numbers aggregate a 28-day (4-week) rolling window; trends compare that window against the 28 days before it.

Categories with fewer than 20 sampled answers in the current window are marked “insufficient data” and publish no rankings — we publish stable numbers, not noise.

recommendation_share = % of sampled answers in the window that recommend the product · avg position = mean rank when answers are ordered lists

§ Statistical robustness

The Index is a sample, not a census — so here are the questions a statistically literate reader should ask before trusting a number on it, answered from the live corpus where the answer is measurable and answered “not yet” where it isn’t.

How many prompts per category?

1–8 prompts per active category (median 7), across 294 categories with a written prompt set. Prompts are neutral, intent-tagged, and stable between samples — see Prompt design.

How many independent runs per category?

Each run is one independent sampled answer: one prompt, one engine, one call, no conversational context carried between them. In the current 28-day window, categories carry 3–112 runs (median 63)294 categories sampled, 271 of them at or above the 20-run publishing gate. Every category page states its own sample size; nothing is published off an undisclosed n.

Which engines, and which models?

Up to 5 engines via their official APIs, and the exact model measured per engine is read live from the collection pipeline and listed in Engines & measured models. Model changeovers are logged and never spliced into one series — the models are part of the measurement, so a change to them is a change to the instrument.

How is variance reported?

Every ranked entry publishes a 95% Wilson score confidence interval on its recommendation share, computed on that entry’s mention-runs over the category’s sample size. Wilson rather than the normal approximation because shares here routinely sit near zero on small samples, where the normal interval leaks below 0 and undercovers.

One honest caveat: the interval treats the window’s runs as independent samples. They aren’t fully — the prompt set is fixed and the same engines are re-sampled — so true uncertainty is somewhat wider than the stated band. Prompt randomization is on the not-yet-measured list on this page.

p = x/n · center = (p + z²/2n) / (1 + z²/n) · half = z·√( p(1−p)/n + z²/4n² ) / (1 + z²/n) · z = 1.96

Category pages render the interval next to the share, so two entries whose intervals overlap can be read as what they are: not distinguishable at this sample size. Each entry also publishes a first-pick rate — the share of its mention-runs where it led the answer’s list — because being named at all and being the top recommendation are different claims. Trends carry a separate comparability gate (see Trend guardrails).

Are prompts randomized or paraphrased?

Not yet. Prompt wording is fixed per category, which buys clean trend comparability — the same question asked over time — at the cost of measuring only that phrasing. A different phrasing of the same buying question can surface a different set of products, and the Index currently cannot tell you how much. Rotating a randomized paraphrase set alongside the fixed set, so phrasing sensitivity is itself measured, is on the roadmap.

Are personas simulated?

No. Prompts carry the audience wording a category’s buyers use (“for freelancers”, “for small teams”), but there are no simulated user profiles, no memory, and no account history behind a run. What an engine recommends to a logged-in user it has learned from is outside what the Index measures.

Is there geographic variation?

Not measured. Runs are issued from a single region without locale hints, so every published number is one geography’s view. Engines can and do localize recommendations; the Index does not currently quantify that, and no number on it should be read as a global average.

What this measurement cannot tell you

AI answers are probabilistic. The same prompt to the same model on the same day can return a different list, which is exactly why nothing here rests on a single answer and why every share ships with an interval instead of a decimal that looks more certain than it is. Engines are also a moving target: vendors change models, retrieval and grounding behavior shift without announcement, and a share can move because the instrument moved rather than because the market did. The guardrails above exist to keep that from being published as a finding.

So: the Index measures what a set of engines recommended, for a fixed set of neutral prompts, in one region, over a 28-day window — not what “AI thinks” and not what any individual user will see. Where a number is too thin to defend, we publish an em-dash instead. The live statistics behind this section are on /data, and if any of it doesn’t hold up we would rather hear it: hello@orbator.io.

§ Model changeovers

Engine vendors retire and replace models. When the measured model behind an engine changes, the Index never splices the two into one series: published aggregation is restricted to the new model’s runs, and trends for that engine are suppressed or flagged until the new model has accumulated a comparable window. The models in the Engines section above are read live and always show what is measured right now.

DateEngineChangeStatus
Jul 21, 2026CLAUDE (Anthropic)claude-haiku-4-5 claude-sonnet-5announced — takes effect when the environment override lands
Jul 21, 2026CHATGPT (OpenAI)gpt-4o-mini chat-latestannounced — takes effect when the environment override lands

§ Entity resolution

Every product mentioned in an answer is extracted and resolved against a canonical registry, so “HubSpot”, “Hubspot CRM” and “hubspot.com” count as one entity. Resolution matches by domain first, then by name and known aliases; cited URLs serve as a confirmation signal.

Generic phrases (“CRM software”, “the tool”) are filtered and never become entities. Newly seen entities and suspected duplicates are auto-flagged for human review; confirmed duplicates are merged with their history re-pointed, and confirmed junk is removed from the rankings.

§ Junk handling & corrections

Word frequency is not a recommendation, so the pipeline enforces rules that keep junk out of published shares:

  • Ambiguous-name corroboration. Products named after dictionary words (ChatBot, Writer, Wave) count only when the sighting is corroborated — the exact cased spelling, a position in a ranked list, or the product’s own domain cited in the answer. “Any chatbot platform” never becomes a sighting of ChatBot.
  • Recommendations, not mentions. A product counts toward recommendation share only when the answer actually recommends it — a ranked-list position or a recommendation cue. Context mentions (“works for Amazon sellers”, “integrates with Salesforce”) are stored but excluded from published shares.
  • Human review & rejection. Newly seen entities and suspected duplicates queue for admin review; confirmed duplicates are merged and confirmed junk is rejected — rejected entities never appear in rankings, and an automated daily integrity check flags any that somehow do.

See something wrong — a misresolved product, a junk entity, a number that doesn’t hold up? Corrections: hello@orbator.io. We would rather pull a number than publish a wrong one.

§ Extraction versions

The rules that turn an answer into counted recommendations — junk filters, source exclusions, corroboration requirements — improve as we find and fix measurement bugs. Every rule change bumps the extraction version, and every stored sighting is stamped with the version that produced it, so analyses can separate data collected under different rules instead of silently mixing them. Current version: 4.

vIn effectWhat changed
1before Jul 9, 2026Original extraction. Generic phrases filtered; no source/platform filters.
2Jul 9 – 21, 2026Review sites, analyst firms, social platforms, and search infrastructure excluded — they are where answers look things up, not recommendations (enforced at the single storage chokepoint from Jul 17).
3Jul 21, 2026AI assistant products (ChatGPT, GitHub Copilot, …) re-admitted as valid entities — only in categories that rank assistant products, where they belong on the leaderboard.
4Jul 21, 2026 — currentDictionary-word product names (ChatBot, Writer, Wave) require corroboration — exact spelling, a ranked list position, or their own domain cited — and context-only mentions ("works for Amazon sellers") no longer count toward recommendation share.

§ Independence

Orbator customers are badged on Index pages for disclosure. Customer status does not affect measurement — the prompts, sampling schedule, extraction, and ranking math are identical for every product, customer or not. Rankings cannot be bought, and no one can pay to be removed.

The Index exists because we sell AI-visibility tooling — that is exactly why the measurement itself has to be, and is, untouched by who pays us.

§ Using the data

Every category is downloadable as CSV from its page and queryable via the free JSON API (/api/index/categories; the measurement parameters on this page at /api/index/methodology; docs at /developers). Free to use with attribution to orbator.io. Questions or corrections: hello@orbator.io.

[ORBATOR]

© 2026 Orbator. All rights reserved.