The AI Recommendation Index measures what AI assistants actually recommend when buyers ask. Every number on an Index page traces back to the process below — and every page discloses its own sample size and window. The measurement parameters on this page (models, gates, guardrails) are read live from the pipeline that collects the data, not hand-maintained copy.
Each category carries a fixed set of buyer-shaped prompts — 6–8 by design; the Statistical robustness section reads the live distribution: the questions real buyers type, like “best [category] for [audience]”, “[category] alternatives”, or “what should I use for [job]”.
Prompts are strictly neutral — they never contain product names, never steer toward or away from any vendor, and stay stable between samples so trends compare like with like.
Categories are measured across up to 5 engines via their official APIs, with web search enabled where the engine supports it. An AI engine is not one thing — answers depend on the exact model behind it, so the Index discloses the model measured per engine:
Each category page shows per-engine shares separately — engines disagree, and that disagreement is part of the data. Not every category runs on every engine; the engines measured are listed per category.
Sampling runs continuously, spread across the batch cycle to avoid time-of-day artifacts. At the current measurement budget, each engine samples each category roughly monthly — cadence scales with budget, and every published number discloses exactly how many sampled answers it rests on.
Published numbers aggregate a 28-day (4-week) rolling window; trends compare that window against the 28 days before it.
Categories with fewer than 20 sampled answers in the current window are marked “insufficient data” and publish no rankings — we publish stable numbers, not noise.
recommendation_share = % of sampled answers in the window that recommend the product · avg position = mean rank when answers are ordered lists
A trend is published only when the two windows are statistically comparable. A trend is suppressed — shown as an em-dash (“not enough comparable data”), never as a direction — when any of these fail:
Trends that pass carry a confidence grade: high, or low (any engine under 20 runs in either window — rendered muted). Share changes within ±1 percentage point are reported as flat; below 50 sampled answers one answer is worth multiple points of share, so the flat band widens to ±5pp.
The Index is a sample, not a census — so here are the questions a statistically literate reader should ask before trusting a number on it, answered from the live corpus where the answer is measurable and answered “not yet” where it isn’t.
1–8 prompts per active category (median 7), across 294 categories with a written prompt set. Prompts are neutral, intent-tagged, and stable between samples — see Prompt design.
Each run is one independent sampled answer: one prompt, one engine, one call, no conversational context carried between them. In the current 28-day window, categories carry 3–112 runs (median 70) — 294 categories sampled, 271 of them at or above the 20-run publishing gate. Every category page states its own sample size; nothing is published off an undisclosed n.
Up to 5 engines via their official APIs, and the exact model measured per engine is read live from the collection pipeline and listed in Engines & measured models. Model changeovers are logged and never spliced into one series — the models are part of the measurement, so a change to them is a change to the instrument.
Every ranked entry publishes a 95% Wilson score confidence interval on its recommendation share, computed on that entry’s mention-runs over the category’s sample size. Wilson rather than the normal approximation because shares here routinely sit near zero on small samples, where the normal interval leaks below 0 and undercovers.
One honest caveat: the interval treats the window’s runs as independent samples. They aren’t fully — the prompt set is fixed and the same engines are re-sampled — so true uncertainty is somewhat wider than the stated band. Prompt randomization is on the not-yet-measured list on this page.
p = x/n · center = (p + z²/2n) / (1 + z²/n) · half = z·√( p(1−p)/n + z²/4n² ) / (1 + z²/n) · z = 1.96
Category pages render the interval next to the share, so two entries whose intervals overlap can be read as what they are: not distinguishable at this sample size. Each entry also publishes a first-pick rate — the share of its mention-runs where it led the answer’s list — because being named at all and being the top recommendation are different claims. Trends carry a separate comparability gate (see Trend guardrails).
Not yet. Prompt wording is fixed per category, which buys clean trend comparability — the same question asked over time — at the cost of measuring only that phrasing. A different phrasing of the same buying question can surface a different set of products, and the Index currently cannot tell you how much. Rotating a randomized paraphrase set alongside the fixed set, so phrasing sensitivity is itself measured, is on the roadmap.
No. Prompts carry the audience wording a category’s buyers use (“for freelancers”, “for small teams”), but there are no simulated user profiles, no memory, and no account history behind a run. What an engine recommends to a logged-in user it has learned from is outside what the Index measures.
Not measured. Runs are issued from a single region without locale hints, so every published number is one geography’s view. Engines can and do localize recommendations; the Index does not currently quantify that, and no number on it should be read as a global average.
AI answers are probabilistic. The same prompt to the same model on the same day can return a different list, which is exactly why nothing here rests on a single answer and why every share ships with an interval instead of a decimal that looks more certain than it is. Engines are also a moving target: vendors change models, retrieval and grounding behavior shift without announcement, and a share can move because the instrument moved rather than because the market did. The guardrails above exist to keep that from being published as a finding.
So: the Index measures what a set of engines recommended, for a fixed set of neutral prompts, in one region, over a 28-day window — not what “AI thinks” and not what any individual user will see. Where a number is too thin to defend, we publish an em-dash instead. The live statistics behind this section are on /data, and if any of it doesn’t hold up we would rather hear it: hello@orbator.io.
Engine vendors retire and replace models. When the measured model behind an engine changes, the Index never splices the two into one series: published aggregation is restricted to the new model’s runs, and trends for that engine are suppressed or flagged until the new model has accumulated a comparable window. The models in the Engines section above are read live and always show what is measured right now.
| Date | Engine | Change | Status |
|---|---|---|---|
| Jul 21, 2026 | CLAUDE (Anthropic) | claude-haiku-4-5 → claude-sonnet-5 | announced — takes effect when the environment override lands |
| Jul 21, 2026 | CHATGPT (OpenAI) | gpt-4o-mini → chat-latest | announced — takes effect when the environment override lands |
Every product mentioned in an answer is extracted and resolved against a canonical registry, so “HubSpot”, “Hubspot CRM” and “hubspot.com” count as one entity. Resolution matches by domain first, then by name and known aliases; cited URLs serve as a confirmation signal.
Generic phrases (“CRM software”, “the tool”) are filtered and never become entities. Newly seen entities and suspected duplicates are auto-flagged for human review; confirmed duplicates are merged with their history re-pointed, and confirmed junk is removed from the rankings.
Word frequency is not a recommendation, so the pipeline enforces rules that keep junk out of published shares:
See something wrong — a misresolved product, a junk entity, a number that doesn’t hold up? Corrections: hello@orbator.io. We would rather pull a number than publish a wrong one.
The rules that turn an answer into counted recommendations — junk filters, source exclusions, corroboration requirements — improve as we find and fix measurement bugs. Every rule change bumps the extraction version, and every stored sighting is stamped with the version that produced it, so analyses can separate data collected under different rules instead of silently mixing them. Current version: 4.
| v | In effect | What changed |
|---|---|---|
| 1 | before Jul 9, 2026 | Original extraction. Generic phrases filtered; no source/platform filters. |
| 2 | Jul 9 – 21, 2026 | Review sites, analyst firms, social platforms, and search infrastructure excluded — they are where answers look things up, not recommendations (enforced at the single storage chokepoint from Jul 17). |
| 3 | Jul 21, 2026 | AI assistant products (ChatGPT, GitHub Copilot, …) re-admitted as valid entities — only in categories that rank assistant products, where they belong on the leaderboard. |
| 4 | Jul 21, 2026 — current | Dictionary-word product names (ChatBot, Writer, Wave) require corroboration — exact spelling, a ranked list position, or their own domain cited — and context-only mentions ("works for Amazon sellers") no longer count toward recommendation share. |
Orbator customers are badged on Index pages for disclosure. Customer status does not affect measurement — the prompts, sampling schedule, extraction, and ranking math are identical for every product, customer or not. Rankings cannot be bought, and no one can pay to be removed.
The Index exists because we sell AI-visibility tooling — that is exactly why the measurement itself has to be, and is, untouched by who pays us.
Every category is downloadable as CSV from its page and queryable via the free JSON API (/api/index/categories; the measurement parameters on this page at /api/index/methodology; docs at /developers). Free to use with attribution to orbator.io. Questions or corrections: hello@orbator.io.
© 2026 Orbator. All rights reserved.