In short: The unit of work is a question somebody would actually type, not a keyword. Ten of them, written in the buyer's words, is enough to run a category and small enough to check by hand every quarter.
Choose them by what an answer decides rather than by volume. There is no reliable volume data for questions asked inside an assistant, and picking by search volume imports a ranking habit into a place where nothing ranks.
Then write them per market rather than translating, because the retrieval behind them does not carry across languages.
Questions, Not Keywords
A keyword is a fragment somebody types into a box that will return a list. A question is what somebody asks when they expect an answer, and it carries the context that a fragment leaves out: who is asking, what they have already ruled out, and what would count as an answer.
That difference is not stylistic. The system rewrites what it receives into its own search terms, so the phrasing you target is not the phrasing that gets searched. What survives the rewrite is the intent and the category, which is what a question carries and a keyword does not.
The test for a good one: could two informed people give different answers to it? "What is a dishwasher drain pump" has one answer and teaches you nothing about your visibility. "How do I know which drain pump fits my dishwasher" has several, and which one an assistant gives is a commercial fact about your category.
Ten, And Why That Number
Ten questions across three or four intents is the smallest set that covers a category and the largest set anyone actually re-runs.
Each question needs three fresh runs per platform to say anything, because a question does not agree with itself. Ten questions across four platforms at three runs is 120 answers to read, which is an afternoon.
Thirty questions is three afternoons, which means it happens once and never again. A baseline nobody repeats is not a baseline.
- 01Category questionsThree or four. No brand names in them. These are the ones that show whether you are a candidate at all.
- 02Comparison questionsTwo or three. Naming an alternative, or asking for one. These show what you are placed beside.
- 03Problem questionsTwo or three. The symptom a buyer describes before they know what to buy.
- 04Brand questionsTwo. Naming you directly, to separate a resolution failure from a classification one.
Choosing By What The Answer Decides
Search keyword selection has a volume number to sort by. Nothing equivalent exists for questions asked inside an assistant, and the tools that publish one are estimating it from search data, which is a different behaviour by a different population.
So sort by consequence instead. For each candidate question, ask what a buyer does immediately after receiving the answer. A question whose answer ends in a purchase decision, a shortlist, or a supplier being ruled out is worth measuring. A question whose answer ends in the buyer reading more is not, however often it is asked.
Where to get them, in order of value: your sales team's inbox, your support tickets, and the questions your own people ask when they join. All three are records of what somebody actually wanted to know, phrased the way they phrased it. Keyword tools come last, and their output has to be rewritten into questions before it is usable.
Per Market, Not Translated
Write the set once in the market's own language rather than translating an English set into it.
A translated question carries English phrasing and English assumptions about what a buyer would ask. The study on this site found the cited sources for a question barely overlap between languages: between 4% and 11%, against 72% to 74% when a question is compared with its own rerun.
A translated set measures how a market answers an English question, which is not a question anyone in that market asked.
This is more work and there is no way around it. It is also the reason a single blended visibility number across markets describes no market that exists.
Write Them Down Somewhere Boring
A spreadsheet, with the question, the market, the intent, and a column per platform per run. Not a tool. The point of a fixed list is that the next measurement asks exactly the same thing, and a list that lives in somebody's head changes between quarters without anybody deciding to change it.
The test plan builder on this site will generate the set and the running order from your category and markets, and you run them yourself.
Key Takeaways
- The unit is a question somebody would ask, not a keyword. The system rewrites your phrasing anyway; what survives is the intent and the category.
- A question worth measuring is one two informed people could answer differently.
- Ten questions across four intents, three fresh runs each. Thirty is a set that gets run once and never repeated.
- Sort by what the answer decides, not by volume. No reliable volume data exists for questions asked inside an assistant.
- Write them natively per market. A translated set measures how a market answers an English question.
Check yourself
Before you move on
Not scored, not recorded, and not part of the certificate. Both answers are settled by a sentence in this lesson, and the reasoning appears whichever option you pick.
- 01
What makes a question worth putting in the measured set?
- 02
Why write the question set natively per market rather than translating one set?