In short: Before any measurement can tell you something changed, you need a number for how much it changes when nothing has. That number is your noise floor and almost nobody measures it.
You get it by asking the same question of the same platform three times in fresh sessions, and recording how much the answers disagree with each other rather than with anything else.
Every result you report afterwards is a comparison against that figure, not against zero, and it has to be measured again rather than assumed to hold.
Why A Control Comes First
A visibility measurement is a comparison, and a comparison against an unstable thing means nothing until you know how unstable it is.
The study on this site exists because of this. Before comparing languages it measured how much a question disagrees with its own rerun, and the answer decided what any language difference was allowed to mean. Without the control, the finding would have been a number with no scale attached to it.
The correction worth knowing about: the first version of that control was wrong by 14 points, because it included repeat pairs where neither run cited anything and those dragged the average down. Fourteen points is larger than most differences reported in this field. A control is not a formality; it is the thing most likely to be quietly wrong.
How To Measure It
Take three questions from your set. For each, on each platform that matters, run it three times in a fresh session, on the same day. Record the sites named in each answer.
Then, for each pair of runs, count how many named sites the two have in common and divide by the total distinct sites across both. That fraction is the agreement between two runs of the same question. Average across pairs and you have the floor.
Two Traps In The Arithmetic
Empty answers. If a run cites nothing, it is not agreeing perfectly with another empty run, and it is not disagreeing completely either. Exclude pairs where neither run cited anything. Including them is the exact error that moved the study's headline by 14 points.
One platform's floor is not another's. A platform that cites twice as many sources per answer has more room to overlap and a different floor. Compute one per platform and never average them into a single number.
The Floor Moves
It is not a constant you measure once. It drifts with model releases, index refreshes and whatever the platform changed without telling anyone.
The study measured self-agreement within a day at 72% to 74%, and across days at a lower figure. That gap is a fact about the system rather than about anyone's website, and it means a measurement taken on Monday and compared with one taken on Thursday carries a wider floor than the same-day number suggests.
Re-measure it when you measure anything. Three questions, three runs, on the same day as the real measurement. It adds twenty minutes to a quarterly run and it is the only thing that keeps the comparison honest as the platforms change underneath it.
What It Is Good For
Two things, and both come up in the same meeting.
It sets the threshold in a goal, so a target is stated as something that clears the noise rather than as a number that sounds ambitious. And it answers the question that ends most reporting arguments: whether the movement in the chart is a result or a rerun.
Key Takeaways
- A control comes before a comparison. Without a noise floor, a visibility number has no scale attached to it.
- Measure it by running the same question three times in fresh sessions and computing how much the runs agree with each other.
- Exclude pairs where neither run cited anything. Including them is what moved a published headline by 14 points.
- Compute one floor per platform. A platform that names more sources has more room to overlap.
- Re-measure the floor whenever you measure anything, on the same day, because it drifts with model releases.
Check yourself
Before you move on
Not scored, not recorded, and not part of the certificate. Both answers are settled by a sentence in this lesson, and the reasoning appears whichever option you pick.
- 01
When computing a noise floor, what must be excluded from the pairs?
- 02
Why compute a separate noise floor for each platform?