Loading...
Loading...
Original measurement
A brand selling into seven European markets does not have one AI visibility problem. It has seven, and the work that fixes one does not carry to the next.
Emad SharakiFull write-up
Ask an assistant a buying question and it names a handful of websites in its answer. Those are the cited domains, and they are the list a brand is trying to be on. This measured how much that list changes when the same question is asked in a different language.
It changes almost completely. Ask the same question twice in English and roughly three quarters of the named sites are the same. Ask it in English and again in Polish and between 4% and 11% are. Of about 750 sites named across the study, more than 675 appeared in exactly one language.
The comparison only means anything against a control, so the first thing measured was how much a question disagrees with itself. That gap is the noise floor, and every language sat at least five times further from English than English sat from its own rerun.
Cross-platform divergence is the finding the field already knows: different assistants cite different sources. This is a second axis, larger and almost unmeasured, inside a single category.
Seven European markets were asked the same eight buyer questions, written natively in each language rather than translated, across four models, three times per cell, in two passes days apart. The cited domain sets were compared using Jaccard similarity.
Correction, published by the author
The first version of this analysis put the noise floor at 0.601. That figure was an artifact: it included 124 repeat pairs where neither run cited anything at all, which dragged the average down.
Corrected, the noise floor is 0.720 to 0.737, and the language separation margin falls from 6.3x to 5.3x. The finding survives. The number moved by 14 points, which is more than most published differences in this field.
Jaccard similarity, two passes. The noise floor is how much a question agrees with its own rerun, so it is the line every other row has to be read against. Anything close to it would be measurement drift rather than a real difference.
| Market | Pass 1 | Pass 2 | Range |
|---|---|---|---|
| Noise floor (self-agreement) | 0.737 | 0.720 | 0.720 to 0.737 |
| Spain | 0.083 | 0.114 | 0.083 to 0.114 |
| Portugal | 0.059 | 0.102 | 0.059 to 0.102 |
| Italy | 0.064 | 0.091 | 0.064 to 0.091 |
| France | 0.049 | 0.084 | 0.049 to 0.084 |
| Germany | 0.039 | 0.072 | 0.039 to 0.072 |
| Poland | 0.039 | 0.072 | 0.039 to 0.072 |
The study set out expecting corpus depth to drive this. It does not. Poland and Portugal are both thin-corpus languages and sit at opposite ends of the table.
| Market | Pass 1 | Pass 2 | Corpus |
|---|---|---|---|
| Poland | 83.9% | 83.8% | Thin |
| Italy | 59.3% | 62.1% | Substantial |
| Germany | 59.6% | 59.9% | Thick |
| France | 54.7% | 56.5% | Substantial |
| Spain | 33.2% | 32.6% | Substantial |
| Portugal | 29.3% | 27.5% | Thin |
| UK | 12.7% | 15.0% | Thick |
The control, and it caught something. Two models drawing on the same third-party index agree with each other more than a question agrees with its own rerun. Two models that search for themselves agree on about 6%.
| Pair | Pass 1 | Pass 2 | Reading |
|---|---|---|---|
| Claude and Gemini (same Exa index) | 0.827 | 0.804 | Above the noise floor: one index, two summarisers |
| GPT and Perplexity (native search) | 0.067 | 0.058 | Genuinely different retrieval |
Carried here with the author's own warning attached: 24 calls per language is too few to read as a pattern. 13% is three calls out of twenty-four, and 33% is eight. Included because the overall rate matters and hiding the weak breakdown would be worse than showing it.
| UK | Germany | France | Spain | Italy | Poland | Portugal |
|---|---|---|---|---|---|---|
| 33% | 25% | 25% | 25% | 25% | 13% | 21% |
The narrative version, with the reasoning behind each decision and the runnable method. This page carries the numbers so they can be quoted without reading it.