In short: A citation is not a link and the difference matters commercially. Being named in an answer builds recognition, being linked sends traffic, and analytics built on sessions can only see the second, so the first is missing from most reporting by construction rather than by choice.
Citation behaviour also differs sharply between platforms, and the overlap between what they cite is far smaller than a single blended visibility score can represent, which makes that score misleading rather than merely imprecise.
Four Different Things People Call A Citation
The word covers four outcomes with very different value, and conflating them is why so much AI visibility reporting is unusable.
| Outcome | What the reader sees | What you get |
|---|---|---|
| Unattributed use | Your information, no source named | Nothing measurable. Your work, someone else's answer. |
| Brand mention | Your name in the text, no link | Recognition and consideration. Invisible in analytics. |
| Named source | Your site listed as a source, not clickable | Authority signal. Still invisible in analytics. |
| Linked citation | A clickable link to your page | Traffic, and the only one your existing reporting can see. |
The reporting trap: If you measure only linked citations, you are measuring the rarest of the four and treating the other three as zero.
A brand named in an answer to a buying question has been recommended by something the reader trusts. That is worth more than most of the impressions you currently count, and your dashboard says it did not happen.
Citation Behaviour Is Platform-Specific
The platforms differ in how many sources they name, what kinds they favour, and how much they weight recency.
In May 2026 the agency 5W Public Relations published an AI Platform Citation Source Index synthesising more than 680 million citations across ChatGPT, Google AI Overviews, Perplexity, Gemini and Claude, drawn from six citation studies conducted between August 2024 and April 2026.
It is an agency synthesis rather than peer-reviewed research, so treat the precise figures as directional. The directions are consistent across the underlying studies, and three of them change what you should do.
Concentration is extreme
In that synthesis the top 15 domains account for roughly 68% of consolidated citation share, and Reddit ranks first across all major engines at around 40% citation frequency. Most of the web is not in the running for most questions.
The uncomfortable implication is that a mention on a heavily cited platform can be worth more than the same content on your own domain. That is not an argument to stop publishing. It is an argument that community presence and earned coverage are GEO work, not a separate discipline that happens elsewhere.
Overlap is small
Only about 11% of domains cited by ChatGPT were also cited by Perplexity in the citations analysed. This is the most consequential number in the field and the least acted upon.
It means winning on one platform tells you close to nothing about your position on another. A single blended "AI visibility score" averages across systems that disagree with each other about which sources exist, which destroys exactly the information you needed.
Freshness weighting differs
In the same synthesis, 36% of Claude's journalism citations came from the previous twelve months, against 56% for ChatGPT. Where a platform weights recency lightly, an article from three years ago is still working for you. Where it weights recency heavily, coverage decays and has to be replenished.
What Makes A Passage Citable
Set the platform differences aside. Across all of them, the same properties make a passage easy to cite, and they are properties of the writing rather than of the domain.
- A claim you can point at. One sentence carrying one fact. A paragraph making three points can be cited for none of them cleanly.
- A visible date. Not just a publish date in the page footer: the date attached to the claim itself, where the claim is time-sensitive. This is what lets a system judge whether your number is still current instead of guessing.
- A stated source. Including when the source is you. "In our analysis of 400 accounts between January and March 2026" is a source. It converts an assertion into something checkable.
- A stable address. A claim that lives at a URL that will still resolve next year is worth more than one behind a rotating carousel or inside a PDF nobody links to.
- An honest scope. Saying what a finding does not cover makes the covered part more usable, because it removes the need to hedge around it.
The rule underneath all five: a citable claim is one a careful editor could repeat without doing additional research. Everything above is a way of removing the reason to check.
The Strongest Position: Be The Source
Every technique above improves your odds of being cited for information that exists elsewhere. There is one position that does not compete on odds at all, and it is being the only place a fact exists.
Original data cannot be summarised away. If you ran the study, published the benchmark, or documented the method, then any discussion of that fact has to name you, because you are part of the fact.
This is why the highest-leverage GEO work is frequently not writing at all: it is measuring something in your industry that nobody has measured and publishing the result properly.
Properly means with the method stated, the sample described, the date attached and the limitations acknowledged. A study that hides its method is not citable, however interesting its conclusion, because nobody can defend repeating it.
Build A Citation Log
You cannot improve citation performance without recording it, and no tool will record it the way your business needs. This takes an hour to set up and ten minutes a month to maintain:
- Choose ten questions where being cited would have commercial value. Buying questions, not brand questions.
- Ask each on at least three platforms, in fresh sessions with no history.
- For every answer, log five fields: the date, the platform, which of the four outcomes above you got, which competitors appeared, and which domains were cited.
- Run each prompt three times rather than once. Record appearances out of three. Answers vary between runs, so a single result is an anecdote.
- Review monthly. Watch two things: your own frequency, and which third-party domains keep appearing. That second list is your earned media target list, and it is more useful than anything you will get from a keyword tool.
Ten questions, three platforms, three runs is ninety data points a month. It is the only measurement in this field that is genuinely yours.
Key Takeaways
- Unattributed use, brand mention, named source and linked citation are four different outcomes. Measuring only links counts the rarest and zeroes the rest.
- Citation share is highly concentrated: roughly 68% across the top 15 domains, with Reddit first at around 40% frequency.
- Around 11% domain overlap between ChatGPT and Perplexity makes blended visibility scores misleading rather than approximate.
- Freshness weighting differs by platform, so earned coverage decays at different rates depending on where you want to be cited.
- Citable passages carry one claim, a date, a source, a stable address and an honest scope.
- Original data is the only position that cannot be summarised away, because it makes your name part of the fact.
Check yourself
Before you move on
Not scored, not recorded, and not part of the certificate. Both answers are settled by a sentence in this lesson, and the reasoning appears whichever option you pick.
- 01
Of the four ways your content can be credited, how many are visible in conventional analytics?
- 02
What is the strongest position a brand can hold in citation terms?