[ Protocol ]
How to measure visibility inside a generative engine
[ In short ]
AI visibility is not a stable ranking. Observe presence, position, sources, and substitutions across a declared set of questions, surfaces, and conditions. Every frequency needs its run count and uncertainty; comparison over time uses the same design and describes an observed change, not proven causality.
- Published
- Reading
- 8 min
Why does the same question give different answers?
Because a generative model samples. Given identical input the phrasing changes, and when the system retrieves documents from the web the retrieved set can change too: it depends on the index at that moment, on internal query rewriting, on the market the request comes from. Add personalisation, account history, and memory, and two people typing the same sentence can read two different answers.
This is why a screenshot is not a measurement. A screenshot proves an answer existed once, for someone, under undeclared conditions. It says nothing about how likely it is to repeat, and it is comparable to nothing.
Variability does not make observation useless. It prevents a single answer from being treated as a stable state. What can be estimated is a frequency under stated conditions, accompanied by the number of runs and its uncertainty. Without those elements, a percentage looks more precise than it is.
What can be measured repeatably?
Four things, and they are worth keeping separate because they answer different questions and are fixed by different work.
- Presence: the brand appears in the answer, yes or no. The rawest figure and the easiest to communicate.
- Position in the answer: first option, one item in a list of alternatives, or a closing aside. The practical effect differs a lot.
- Retrieved sources: which URLs the answer used. The most operational figure, because it points at where to act.
- Substitutions: who gets recommended when you do not appear. The figure that makes the measurement discussable in a commercial meeting.
How do you build the question set?
Not from keywords. Keywords are fragments; what you need are complete questions, written the way a person mid-decision would write them. The practical difference is that a keyword produces a generic answer, while a question with context produces the answer your customer will actually see.
The set must be built by intent, covering the whole decision path. Discovery: who solves this problem. Comparison: which of these two. Vendor selection: who do I go to with this constraint. Objection: what are the limits of this approach. Verification: is this company reliable. Each intent produces structurally different answers, and a set covering only discovery tells a fifth of the story.
Then market and language must be declared, because they shift the result more than almost anything else. The same question in Italian and in English does not retrieve the same sources. Measuring in English a business that sells in Italy produces an elegant, useless number.
There is no universal number of questions. Set size depends on the quantity being estimated, the intents and markets to cover, observed variability, and the cost of repetition. Questions should be stratified by intent and sized around the future comparison: a broad baseline that cannot be repeated is not a useful measurement design.
How many runs per question?
More than one, using the same design for rows you intend to compare. A single run is one observation. Repetition supports a frequency estimate, but runs do not automatically become independent: model behavior, indexes, query fan-out, and retrieved documents can introduce dependencies the report must acknowledge.
Five runs may support exploration, not a precise probability estimate. Run count should follow the required resolution, initial variability, and the smallest difference that would change a decision. Binomial or Bayesian intervals make uncertainty visible; more runs narrow it rather than removing it.
Session conditions must be defined and repeated. A session without memory reduces some sources of variation but does not simulate a personalized user experience. When the case requires it, keep a controlled sample separate from a realistic sample instead of mixing their results.
What gets recorded on each run?
The minimum needed for someone else to reconstruct the row six months later. If a field is missing, that row is not comparable and should be dropped rather than interpreted.
| Field | Why it matters |
|---|---|
| Verbatim question | Lets the measurement be repeated word for word |
| Engine and date | Features and indexes change over time |
| Surface and mode | Consumer UI, API, web search, and conversational modes are not equivalent |
| Available version or identifier | Reduces ambiguity as models and products change |
| Market and language | Shifts retrieved sources more than any other variable |
| Time and session condition | Makes context, memory, and possible temporal effects explicit |
| Run number | Separates frequency from a single accident |
| Brand cited, yes or no | The presence figure |
| Position in the answer | First option, alternative, marginal aside |
| Retrieved sources | Points at what to work on, and the most neglected field |
| Competitors cited | Shows who holds the space when you do not |
| Full answer text | Lets you reread the context without rerunning the test |
How do you compare a new measurement against the baseline?
By using the same declared protocol and documenting what cannot remain fixed. Questions, run count, market, language and session conditions stay comparable where possible; engine and surface updates are recorded. If a question's wording changes, that row leaves the historical comparison and becomes a separate exploratory sample.
The comparison returns categories that are better communicated separately than compressed into a single index. Gains: questions where you did not appear and now do. Losses: the reverse, which should be investigated without assigning an automatic technical or competitive cause. New entrants: competitors who were not there and now hold the observed space.
On the link between action and result, honesty is required. Between a baseline and a re-measurement many things change that you do not control: models get updated, indexes move, competitors publish. The comparison shows the observed position moved, in the expected direction or not. It does not prove it moved because of you. Presenting the second as proven is overreach, and eventually someone checks.
How much does a figure have to move to count?
This is the question missing from almost every report. Comparing gains and losses means little unless the protocol declares which magnitude counts as a signal, because without that threshold every fluctuation can become a convenient story.
The starting point is arithmetic. With five runs per question the resolution is one fifth: the smallest observable change is twenty points. This does not create a universal threshold. Moving from two out of five to three out of five remains compatible with substantial uncertainty and should be described as exploratory evidence.
Before collecting data, declare the smallest difference that would change a decision and choose a design capable of detecting it. For one question, report counts and intervals. At set level, preserve intent strata and use a paired comparison or bootstrap where appropriate, without assuming that intentionally selected questions represent a population.
The reason the threshold must be declared in advance is less technical and more uncomfortable: deciding what counts after seeing the numbers is how you always find what you hoped to find. Writing it into the protocol before the first cycle costs nothing and makes the second cycle arguable by anyone.
A question oscillating between two out of five and three out of five does not establish improvement or decline. It documents instability under the observed conditions and indicates that more data, a different period, or a better-defined question is needed.
How often should you re-measure?
Ninety days can be a practical interval when change requires publishing, indexing, and distribution. It is not a universal threshold: market, engine, intervention type, and set size all change what can be observed. Re-measuring sooner makes sense only when a declared condition has genuinely changed.
The cost of re-measuring is predictable and worth knowing before you size the set: it grows as questions times runs times engines. Sixty questions, five runs, four engines is twelve hundred queries. Double any one of the three and you double the cost of every future cycle, not just the first.
The practical consequence is that a larger set is not automatically more accurate. If it reduces repetition, breaks the strata, or prevents a second cycle, it may produce more rows but less decision value. Size the set around the budget for the whole design, not the volume of the first report.
- Set the re-measurement interval before the work, based on the observed change and the conditions of the set.
- Off cycle, only when something substantial changes: a launch, a repositioning, publishing content aimed at a specific question.
- Size the set on the cost of the second cycle, not the first. The second is where the value comes from.
- When budget forces a reduction, protect the essential intents and comparability. The trade-off between questions, runs, and engines must be justified rather than automated.
How do you challenge a result?
This is the question that separates an analysis from an assertion, and it is worth asking anyone selling you a measurement. If the answer is that the figure cannot be verified, what you bought is not a measurement.
A challengeable result needs a few things, all of them in the report: the question written word for word, the engine, the date, the market, the language, and the number of runs. Those fields let someone repeat the declared protocol and assess the new observation while recognising that engines and retrieval may have changed.
The criterion for judging the challenge is systematicity, not the single run. Rerun the question once and get a different result and you have disproved nothing: you have observed the variability the method already assumes. Get it repeatedly and in the same direction, and a condition has changed or was badly declared, and either way the row needs correcting.
Whoever produces the analysis should want those challenges. They are how a baseline stays reliable over time instead of becoming a document nobody checks any more.
What does this measurement not prove?
It does not prove revenue. It measures presence in an answer, which sits far upstream of a sale. Connecting the two requires other data, and presenting presence as revenue is the fastest way to lose credibility at the first check.
It does not prove causality, for the reason above. It shows movement under stated conditions.
It does not prove the real user sees the same thing. Clean sessions serve reproducibility, not simulation of a personalised experience. That is a deliberate choice: a comparable number is preferred over a realistic but unrepeatable one.
It is not a statistically representative sample. The question set is chosen on purpose, to cover the intents that matter commercially, not drawn at random from a population. That makes it useful for deciding and unsuitable for presenting as an estimate of a universe.
It does not prove total coverage. It measures the questions you chose. If the set is built badly, the result is internally consistent and useless at the same time. The quality of the question set is the most fragile point in the whole process, and the first thing I would ask to see if I were buying an analysis.
[ What to take away ]
- Repeat every question several times and record a frequency. A single run is an observation, not a trend.
- Always declare market, language, engines, date, and session conditions. Without those fields the comparison does not exist.
- Record the retrieved sources. That is the field telling you where to act, and the one missing from almost every report.
- Communicate gains, losses, and new entrants separately. A single index hides exactly the useful information.
- Do not attribute movements to your own work. Show the movement and state the limits.
[ Related service ]
[ comparison over time ]
Explore AI Visibility Monitor[ Related notes ]
[ Sources ]
How is your brand being read?
Tell me about the market, the questions that matter, and what you cannot observe today. Within 48 hours, I will indicate whether you need a baseline, a technical intervention, or a comparison over time.