There is a new category of software that tells you whether ChatGPT recommends your business. We build one. So does everyone else, suddenly.
Here is the uncomfortable part: almost none of it can survive a serious question about how it works.
The critique is already out there, and it is correct. These tools test questions nobody actually asks. They count any time your name appears as a win. They report a single number with no sample size, no margin of error, and no way to check it.
We could argue with that. Instead we rebuilt our own measurement, then pointed it at ourselves. This is what it found.
34.7% of the time an AI names a product, it is not recommending it
Every time one of the five engines names a product in an answer, we now sort what the answer actually did with it. Across 29,511 sightings covering 4,162 products in 105 categories, here is the split:
| What the answer actually did | Share |
|---|---|
| Listed it among options | 38.9% |
| Mentioned it in passing | 33.9% |
| Explicitly recommended it | 26.4% |
| Warned against it | 0.8% |
Only about a quarter of the time does an AI actually tell a buyer to use the thing. A third of the time it is scenery — the product appears because it integrates with something, or as a comparison, or as context.
Under mention-counting, all 29,511 of those count the same. They are all "visibility."
The 246 that should end the argument
Two hundred and forty-six of those sightings are answers actively steering buyers away — avoid this one, I would not use this, stay away. They involve 176 different products across 29 categories.
Under mention-counting, every one of those scored as a win for the product being warned about.
That is not a rounding error in a metric. It is the metric pointing the wrong way. A product could sit near the top of a visibility ranking specifically because AI keeps telling people not to buy it.
We know because we did it. Our own Index counted mentions, and a product sat high in it while the answers said avoid.
What we changed
We stopped counting mentions. The headline number is now a recommendation rate: how often an AI actually told a buyer to use you. Warnings are excluded everywhere downstream — they cannot be quietly rounded into visibility. The old number still exists, labelled as what it is: a count of times your name came up.
Every number carries its own uncertainty. Shares publish with a 95% confidence interval and the sample behind them. Forty-two percent from two hundred answers and forty-two percent from twenty answers are not the same claim, and we were printing them the same way. A category thin enough that the interval would be meaningless now publishes nothing at all — the floor is 20 answers, and no published category sits below it.
We wrote down how it works, and what it cannot do. The methodology page reads live from the running system, so it cannot describe a design we are not actually using. It has a section titled what we cannot see. Every rule change is dated in a public changelog, and our build fails if the rules change without an entry.
We started keeping the record. Every published ranking is frozen daily into a dated, append-only archive — currently 1,355 snapshots across 271 categories. Nothing is ever edited; a correction appends a new version and says what changed. Ask what we published on a date and you get the record, not a recomputation. If we have no record for that day, the answer is "no record."
What we have not solved
We cannot see inside anyone's actual conversation with ChatGPT. Nobody outside OpenAI can. What people ask is shaped by memory, personalisation, and whatever they said three messages earlier, and no tool on the market observes that — including ours. Anyone claiming otherwise is selling you a model's guess with a confident face on it.
So we do not claim to know what any individual buyer was told. We claim something narrower and checkable: for a stated question, on a stated date, against stated models, here is what the answers said, here is how many we read, and here is how sure we are.
These numbers are a snapshot, not a trend. The four-way classification has been running for days, not months, and the archive started this month. Some categories are thinner than we would like. We would rather tell you that than round it off.
The standard
Ask any vendor in this category four questions.
1. Where do your prompts come from?
2. Does a mention count the same as a recommendation?
3. What is the margin of error on that number?
4. What did you publish last month, and can you show me?
We think you should ask us first. The methodology is public, the changelog is dated, the archive is on the record, and the limitations are written down by us rather than discovered by you.
An AI recommendation is becoming the thing that decides which businesses get found. The measurement of it should be boring, checkable, and occasionally embarrassing to the people doing the measuring.
Ours was. We fixed it in public.
---
Method. 29,511 entity sightings from answers collected across ChatGPT, Claude, Gemini, Perplexity and Grok, covering 4,162 products in 105 categories, under extraction version 7 (the four-way classification). Classification is deterministic cue-matching first, with a single model call only for genuinely ambiguous mentions. Per-product samples are still small, which is why no individual company is named here — that would fail the 20-answer floor we hold our own published categories to. Full method: orbator.io/ai-index/methodology