How-To Guide · Updated August 2026

How to Find Out Which Sources AI Engines Read for Your Category

When an AI assistant recommends products in your category, it is not consulting a ranking — it is reading a specific, finite set of pages and repeating what they collectively say. Finding that set is the single highest-leverage move in AI visibility, because it turns “get famous on the internet” into a to-do list. Most tools in this market will tell you whether you appear in answers. This guide covers the harder question: which pages decided the answer — with the full worked example from a category we resolved end to end, so you can judge the method instead of taking our word.

Can you actually see what AI reads before it recommends?

Yes — two of the five major engines show sources on every answer, and the rest can be sampled. The honest method has four steps: ask the buyer's question many times, not once; collect every citation the engines expose; resolve each cited domain to its real owner; and only then read the list. The owner-resolution step is the one almost nobody does, and it is where the surprises live — when we ran the full pipeline on legal software, all 22 of the most-cited pages turned out to be owned by vendors competing in the same answer. The first genuinely independent source was a Reddit thread, at rank 23. If you want the same view of your own category, the free check shows the sources each engine read, per answer.

Step 1. Ask the buyer’s question — enough times to mean something

One answer is an anecdote. The engines vary run to run, so sources only become a list you can act on when you sample.

Ask the way a real buyer asks, in their words, and ask repeatedly. AI answers are not deterministic: the same question re-asked minutes later can cite different pages and name different products. A single screenshot tells you what one roll of the dice read. A sample tells you which sources keep showing up — and those recurring sources are the ones actually steering your category.

For scale: our published legal-software audit sampled 125 answers across ChatGPT, Claude, Gemini, Perplexity and Grok over 28 days, which produced 1,188 resolvable citations — roughly nine or ten cited sources per answer. You do not need that volume to start, but you do need more than one run per engine, and you should record the date, because answers drift as engines re-crawl ( every figure we publish carries its window for exactly this reason — methodology).

Step 2. Collect the citations the engines expose

Perplexity and Gemini show sources on every answer; ChatGPT shows them when it browses. Capture the URLs, not just the domains.

Perplexity numbers its sources on every answer, which makes it the most transparent engine to start with. Gemini links its grounding sources. ChatGPT cites when it searches the web, and so does Grok. Claude cites when browsing is on. Capture the full URLs: two different pages on the same domain can play completely different roles — a vendor’s pricing page and that same vendor’s “best tools like us” listicle are different animals, and the second one is the one steering answers.

Keep the citation attached to its answer. You want to know not just that a page was read, but in how many answers it was read — a page cited once is noise, a page cited across a quarter of your sampled answers is infrastructure.

Step 3. Resolve every cited domain to its real owner

The step almost nobody does, and the one that changes the story: who owns the page that is steering the answer?

For each cited domain, ask one question: is this source independent of the products it discusses, or does it belong to one of them? Corporate registrations, acquisitions and white-label content make this genuinely laborious — sites that read like neutral publications are routinely owned by a competitor in the ranking they publish. There is no shortcut here; it is lookup work. It is also where the real picture appears.

When we resolved all 1,188 citations in the legal-software category, at most 32.9% of sources were independent after human verification — at least two of every three pages the engines read belonged to the vendors' own web estate. Across our whole measured corpus the pattern holds: 48.8% of citations in AI answers point at a vendor competing in that same category (90,960 of 186,331 citations, 28-day window). The reading list behind “neutral” AI advice is, to a first approximation, the industry writing about itself. Live figures on /data.

Step 4. Read the list like a to-do list

Sort resolved sources into three piles: places you can join, pages you can earn, and questions nobody has answered yet.

Joinable surfaces first: directories, review platforms and databases that appear in your citation list are fast, cheap and under your control — a complete, well-described profile hands the engines the exact sentences they repeat. Earnable pages second: the editorial articles and rankings your category’s answers lean on. And third, the most underrated pile: buyer questions whose answers cite weak or missing pages. Those are open slots, and a clear page on your own domain that answers the question in the asker’s words is precisely the kind of source engines pick up — in consumer categories we measure, the engines’ top citations are routinely product sites that simply answered the question plainly.

One warning from our own data: presence is not endorsement. 34.7% of the time an AI names a product, it is not recommending it — it is listing it, mentioning it in passing, or occasionally steering the buyer away (measured across 29,511 sightings; the full split is in the measurement-problem study). When you re-check your category after doing the work, score what the answer did, not whether your name occurred.

FAQ

Do the engines read the same sources as each other?

Only partly, and that is why single-engine checks mislead. Some sources in a category are read by four of the five engines we track; others are one engine’s habit. The overlap set is the highest-value real estate, and the engines genuinely disagree about the final answer far more often than people expect — the live agreement rate is published on our /data page.

Can I just ask the AI which sources it used?

When the engine shows citations, believe the citations. When it does not, its self-report is unreliable — models routinely produce plausible-sounding source lists that do not match what was actually retrieved. Use the exposed citations, sampled across runs, and skip the interrogation.

How often does the reading list change?

Continuously at the edges, and abruptly when engines ship ranking changes. That is why a source list should carry its measurement window, and why we re-measure weekly rather than treating any snapshot as permanent.

Is there a free way to see the sources for my category?

Yes. Orbator’s free check asks your category’s buyer questions across ChatGPT, Claude, Gemini, Perplexity and Grok and shows the sources each engine read, per answer, no card required. The paid work — full sampling, owner resolution, the complete category read — is what our published audits do, and one is public so you can inspect the method end to end.

Numbers on this page: Orbator Category Influence Audit #1 — AI tools for lawyers, v1.2 (125 sampled answers, 5 engines, 1,188 resolved citations, 28-day window ending 2026-08-09; independence at most 32.9% after human verification; public PDF at /audit-1-ai-tools-for-lawyers.pdf) · global corpus vendor-owned citation share 48.8% (90,960 of 186,331, 28-day window, pulled 2026-08-17) · mention-vs-recommendation split from the measurement-problem study (29,511 sightings, 105 categories). Every figure carries its n and window; the live values are on /data.

The answer to every AI recommendation has a reading list. Find yours.

See Your Category’s Sources Free →
[ORBATOR]

© 2026 Orbator. All rights reserved.