In short: The major AI assistants do not share one index or one set of rules. They differ in whether they search the live web at all, how many sources they cite per answer, which kinds of sources they favour, and whether a citation is a clickable link or a passing mention.
One large 2026 synthesis found that only about 11% of domains cited by ChatGPT were also cited by Perplexity. Treating "AI" as a single destination is the most common and most expensive mistake in GEO.
Two Ways AI Finds Information
Every answer an AI assistant gives is assembled from one of two sources, and often both. Knowing which one you are competing for changes what you do.
Training the model already knows
Every model is trained on an enormous body of text: web pages, books, code, forums. What it absorbs becomes "parametric knowledge", stored in the model's weights rather than looked up.
- Always available, with no lookup and no latency
- Frozen at a cutoff date, so recent events are missing
- Cannot be corrected without retraining or fine-tuning
- Carries no citation, because there is no document to point at
This is the layer that decides whether a model has heard of your brand at all. You cannot optimise it directly and you cannot rush it. It is built by being written about, repeatedly, in sources the model was trained on.
Retrieval: what the model looks up
When a question needs current or specific information, the assistant issues its own searches, reads the results, and writes an answer from them. This is retrieval-augmented generation, and it is where citations come from.
- Reflects the live web, so new content can appear within days
- Produces named sources, sometimes as links
- Depends entirely on your pages being crawlable and readable
- Is the layer you can actually influence this quarter
Why the distinction matters: If a model recommends a competitor from training data, no amount of on-page work moves it this month. If it recommends them from retrieval, the gap is a content and crawlability problem you can close.
The Platforms Differ More Than They Resemble Each Other
It is tempting to optimise for "AI search" as one thing. The published evidence says that is wrong.
In May 2026, the agency 5W Public Relations published an AI Platform Citation Source Index synthesising more than 680 million individual citations across ChatGPT, Google AI Overviews, Perplexity, Gemini and Claude, drawn from six large citation studies conducted between August 2024 and April 2026.
It is an agency synthesis rather than peer-reviewed research, so treat the precise figures as directional. The direction, however, is consistent across every study in it.
They cite different amounts
Perplexity cites far more sources per answer than its peers, roughly twice as many URLs as ChatGPT and closer to three times as many as Gemini.
Perplexity is built to show its working; ChatGPT is built to give you an answer. That single design difference changes your odds: on Perplexity you are competing for a place in a list, on ChatGPT you are competing to be the source.
They favour different kinds of sources
- ChatGPT leans heavily on reference and community sources. Wikipedia alone has accounted for between a quarter and roughly half of its top-10 citation share in published samples, with Reddit close behind.
- Perplexity favours primary sources and academic databases, rewarding original documents over commentary about them.
- Claude leans toward established journalism, and toward older journalism: in the same synthesis only 36% of its journalism citations came from the previous twelve months, against 56% for ChatGPT.
Reddit ranks first across all major engines at roughly 40% citation frequency, and the top 15 domains capture around 68% of consolidated citation share. Concentration is high. Most of the web is not in the running.
They barely overlap
The most useful number in the whole dataset is the smallest. Across the citations analysed, only about 11% of domains cited by ChatGPT were also cited by Perplexity. Winning on one platform tells you almost nothing about your position on another.
What This Means For Your Content
| Platform | Rewards | Practical implication |
|---|---|---|
| ChatGPT | Reference-grade and community sources | Entity clarity and third-party mentions matter more than your own page copy |
| Perplexity | Primary sources, original documents | Publish the data, not a summary of someone else's |
| Claude | Established journalism, less freshness-driven | Earned coverage outlasts publishing cadence |
| Google AI Overviews | Existing search visibility | Traditional SEO still carries most of the weight |
| Gemini | Fewer, more selective sources | Being second best is often being invisible |
Three conclusions follow, and they are more useful than any tactic list:
- Pick your platform before you pick your tactics. Your buyers are not spread evenly across these tools. Optimising for all five at once means optimising for none.
- Original material outperforms good summaries. Every platform above rewards being the document rather than the description of it, and Perplexity does so explicitly.
- Measure per platform or do not claim to measure. With roughly 11% domain overlap, a single "AI visibility score" averages away the only information worth having.
How To Check This Yourself
Published studies age quickly, and every vendor publishing one has an interest in its conclusion. Run your own check rather than inheriting someone else's:
- Write down ten questions a real buyer would ask, in their words, not yours.
- Ask each of them on at least three assistants, in a fresh session with no history.
- Record, for each answer: are you mentioned, are you cited, is the citation a live link, and which competitors appear alongside you.
- Repeat monthly. The absolute numbers matter far less than the direction they move.
Ten questions across three platforms is 30 data points and about an hour of work. That is a small enough exercise to actually do, and it will tell you more about your position than any published index.
A note on the 40% figure you will see quoted: The original GEO research (Aggarwal et al., arXiv:2311.09735, submitted November 2023, later published at KDD 2024) reported that its methods boosted visibility by up to 40%.
That result was measured on GEO-bench, the authors' own benchmark, under their own conditions. It is a real finding and a reasonable reason to take GEO seriously. It is not a promise about your site, and anyone quoting it as one has not read the paper.
Key Takeaways
- AI answers come from training data or from retrieval. Only the second responds to work you do this quarter.
- The platforms are not variations on one system. They differ in citation volume, source preference and freshness bias.
- The 2026 citation synthesis found roughly 11% domain overlap between ChatGPT and Perplexity, which makes platform-specific measurement not optional.
- In the same synthesis, citation share is highly concentrated: around 68% sits with the top 15 domains.
- Run your own ten-question check before trusting any published visibility figure, including the ones on this page.
Check yourself
Before you move on
Not scored, not recorded, and not part of the certificate. Both answers are settled by a sentence in this lesson, and the reasoning appears whichever option you pick.
- 01
Roughly what share of domains cited by ChatGPT were also cited by Perplexity, in the 2026 synthesis this lesson cites?
- 02
Why is Perplexity the most responsive surface for early testing?