Skip to content

[ Definition ]

How to measure visibility inside a generative engine

[ In short ]

You do not measure a ranking, you measure a frequency. Fix a question set, run each question several times in clean sessions, and record how often the brand appears and which sources the answer used. The useful number is frequency under stated conditions, compared against the same measurement repeated later.

Published
Reading
8 min

Why does the same question give different answers?

Because a generative model samples. Given identical input the phrasing changes, and when the system retrieves documents from the web the retrieved set can change too: it depends on the index at that moment, on internal query rewriting, on the market the request comes from. Add personalisation, account history, and memory, and two people typing the same sentence can read two different answers.

This is why a screenshot is not a measurement. A screenshot proves an answer existed once, for someone, under undeclared conditions. It says nothing about how likely it is to repeat, and it is comparable to nothing.

Variability does not make measurement impossible though. It makes point measurement impossible. What stays measurable is frequency: across twenty runs of the same question, how many times the brand appears. That is a number you can repeat, compare, and argue about.

What can be measured repeatably?

Four things, and they are worth keeping separate because they answer different questions and are fixed by different work.

  • Presence: the brand appears in the answer, yes or no. The rawest figure and the easiest to communicate.
  • Position in the answer: first option, one item in a list of alternatives, or a closing aside. The practical effect differs a lot.
  • Retrieved sources: which URLs the answer used. The most operational figure, because it points at where to act.
  • Substitutions: who gets recommended when you do not appear. The figure that makes the measurement discussable in a commercial meeting.

How do you build the question set?

Not from keywords. Keywords are fragments; what you need are complete questions, written the way a person mid-decision would write them. The practical difference is that a keyword produces a generic answer, while a question with context produces the answer your customer will actually see.

The set must be built by intent, covering the whole decision path. Discovery: who solves this problem. Comparison: which of these two. Vendor selection: who do I go to with this constraint. Objection: what are the limits of this approach. Verification: is this company reliable. Each intent produces structurally different answers, and a set covering only discovery tells a fifth of the story.

Then market and language must be declared, because they shift the result more than almost anything else. The same question in Italian and in English does not retrieve the same sources. Measuring in English a business that sells in Italy produces an elegant, useless number.

A reasonable size for a first baseline sits between forty and sixty questions. Below twenty the sample is too small to tell a signal from an accident. Above a hundred the cost of re-measuring becomes the reason the baseline never gets repeated, and that is the most expensive mistake of all.

How many runs per question?

More than one, and always the same number for all of them. A single run records one sample of a stochastic process and treats it as a fact. Repeating the same question instead yields a frequency, which is the only number that survives a comparison over time.

With five runs per question you can already separate stable behaviour (appears five out of five, or zero out of five) from unstable behaviour. With ten you start reading the middle bands with some confidence. The choice is not a statistical law, it is a declared trade-off between cost and resolution, and it belongs in the report next to the number.

Runs must happen in clean sessions: no history, no active memory, no linked profile, no prior question in the same thread. Not because that is the real user's condition, which is personalised, but because it is the only reproducible one. A measurement inside an account with three years of history measures the account, not the brand.

What gets recorded on each run?

The minimum needed for someone else to reconstruct the row six months later. If a field is missing, that row is not comparable and should be dropped rather than interpreted.

Nine fields. Eight are descriptive; only one (position) needs a judgement criterion, and that criterion must be written down.
FieldWhy it matters
Verbatim questionLets the measurement be repeated word for word
Engine and dateFeatures and indexes change over time
Market and languageShifts retrieved sources more than any other variable
Run numberSeparates frequency from a single accident
Brand cited, yes or noThe presence figure
Position in the answerFirst option, alternative, marginal aside
Retrieved sourcesPoints at what to work on, and the most neglected field
Competitors citedShows who holds the space when you do not
Full answer textLets you reread the context without rerunning the test

How do you compare a new measurement against the baseline?

By holding every variable fixed except time. Same questions word for word, same number of runs, same engines, same market, same language, same session conditions. Change even the phrasing of one question and that row leaves the comparison: it is not a worse figure, it is a different figure.

The comparison returns three categories, and they are better communicated separately than compressed into a single index. Gains: questions where you did not appear and now do. Losses: the reverse, and worth reading first because they often signal a technical problem rather than a competitive one. New entrants: competitors who were not there and now hold the space.

On the link between action and result, honesty is required. Between a baseline and a re-measurement many things change that you do not control: models get updated, indexes move, competitors publish. The comparison shows the observed position moved, in the expected direction or not. It does not prove it moved because of you. Presenting the second as proven is overreach, and eventually someone checks.

How much does a figure have to move to count?

This is the question missing from almost every report, including the way I had written it up to here. Saying to compare gains and losses is worth little without declaring which magnitude counts as a signal, because absent that threshold every wobble becomes a story.

The starting point is arithmetic. With five runs per question the resolution is one fifth: the smallest observable change on that question is twenty points. So moving from two out of five to three out of five is a single run that went differently, which is exactly the magnitude of noise the method already assumes. Treating it as an improvement is reading measurement error.

Hence two rules I use, and they have to be fixed before looking at the data rather than after. At single-question level I only read movements that touch the extremes: from zero out of five to three or more, or the reverse. Those are the only ones wide enough not to be explained by one run. At set level I look at the aggregate instead, where the noise of individual questions partly cancels, and where a few percentage points across forty questions is more reliable than a dramatic move on one.

The reason the threshold must be declared in advance is less technical and more uncomfortable: deciding what counts after seeing the numbers is how you always find what you hoped to find. Writing it into the protocol before the first cycle costs nothing and makes the second cycle arguable by anyone.

Finally, a question oscillating between two out of five and three out of five at every reading is not a failure of the measurement. It is information: on that question your presence is unstable, and documented instability is a different result from absence.

How often should you re-measure?

Ninety days is the interval that holds up best. Below that threshold the effects of publishing are rarely legible: content has to be indexed, retrieved, and picked, and that chain does not close in two weeks. Re-measuring monthly almost always produces noise that will be read as signal, and that is the most efficient way to make wrong decisions from real data.

The cost of re-measuring is predictable and worth knowing before you size the set: it grows as questions times runs times engines. Sixty questions, five runs, four engines is twelve hundred queries. Double any one of the three and you double the cost of every future cycle, not just the first.

Hence the counterintuitive consequence: an oversized set is not more accurate, it is less useful. It becomes the reason the second measurement never happens, and a baseline with no comparison is a document, not a measuring system. Forty questions repeated four times a year beats a hundred and fifty measured once.

  • Ninety days for a full re-measurement of the same set, under the same conditions.
  • Off cycle, only when something substantial changes: a launch, a repositioning, publishing content aimed at a specific question.
  • Size the set on the cost of the second cycle, not the first. The second is where the value comes from.
  • If the budget allows only one axis, cut the number of questions and keep the runs. Fewer questions measured well beats more questions measured once.

How do you challenge a result?

This is the question that separates an analysis from an assertion, and it is worth asking anyone selling you a measurement. If the answer is that the figure cannot be verified, what you bought is not a measurement.

A challengeable result needs a few things, all of them in the report: the question written word for word, the engine, the date, the market, the language, and the number of runs. With those fields anyone can rerun the same question in a clean session, under the same conditions, and compare.

The criterion for judging the challenge is systematicity, not the single run. Rerun the question once and get a different result and you have disproved nothing: you have observed the variability the method already assumes. Get it repeatedly and in the same direction, and a condition has changed or was badly declared, and either way the row needs correcting.

Whoever produces the analysis should want those challenges. They are how a baseline stays reliable over time instead of becoming a document nobody checks any more.

What does this measurement not prove?

It does not prove revenue. It measures presence in an answer, which sits far upstream of a sale. Connecting the two requires other data, and presenting presence as revenue is the fastest way to lose credibility at the first check.

It does not prove causality, for the reason above. It shows movement under stated conditions.

It does not prove the real user sees the same thing. Clean sessions serve reproducibility, not simulation of a personalised experience. That is a deliberate choice: a comparable number is preferred over a realistic but unrepeatable one.

It is not a statistically representative sample. The question set is chosen on purpose, to cover the intents that matter commercially, not drawn at random from a population. That makes it useful for deciding and unsuitable for presenting as an estimate of a universe.

It does not prove total coverage. It measures the questions you chose. If the set is built badly, the result is internally consistent and useless at the same time. The quality of the question set is the most fragile point in the whole process, and the first thing I would ask to see if I were buying an analysis.

[ What to take away ]

  • Repeat every question several times and record a frequency. A single run is not a measurement.
  • Always declare market, language, engines, date, and session conditions. Without those fields the comparison does not exist.
  • Record the retrieved sources. That is the field telling you where to act, and the one missing from almost every report.
  • Communicate gains, losses, and new entrants separately. A single index hides exactly the useful information.
  • Do not attribute movements to your own work. Show the movement and state the limits.

[ Author ]

Nicola Dussin

Founder of Creaitivo. Every analysis is run directly by me.

LinkedIn profile

Want the same reading on your case?

The automated preview gives a first signal in seconds. I prepare the useful reading myself, and it arrives within 48 hours.

Send me your case
All insights