TL;DR
The field agreed on the right sentence and drew the wrong conclusion. The standard response to “AI visibility is a distribution” is to sample more and publish a confidence interval, which treats the spread as measurement error around a true score.
It is not error. Inclusion in an AI answer is binary, so at the level of a single prompt the variance is a fixed function of the mean and carries no information the mean does not already hold.
All the informative spread sits between prompts. A typical basket is not one coin but a mixture of prompts you always win, prompts you sometimes win, and prompts you never enter. A mean computed across that mixture describes none of them.
Three instruments follow — the Visibility Shape, the Dispersion Check and the Decision Functional — plus why the shape matters more to a link building budget than the score does.
The field agreed on the right sentence, then drew the wrong conclusion
By mid-2026, “visibility is a distribution, not a number” had become the standard caveat in every serious piece of writing about brand presence in AI answers. It is a true sentence. What has been built on top of it is not.
The remedy attached is always the same: run each prompt several times, average, publish an interval. Vendors now advertise run counts the way they once advertised keyword databases. But look at what that assumes. Averaging is correct under one specific model of the world — that a true score exists and what you observe is that score plus noise. That model licenses exactly one operation: average the noise away until the signal shows. Every confidence interval you have seen in an AI visibility report is an artefact of someone assuming it without saying so.
Answer engines do not work that way, and the evidence sits in the same papers everyone cites for the run-count recommendation. The variation is not error on top of a true score. A large part of it is the thing being measured — and when the spread is the phenomenon rather than the interference, sampling harder buys a more precise estimate of a quantity that describes nothing in particular.
One scoping note. This piece is about the shape of the distribution at a moment in time — what kind of object your visibility score summarises. Why the number moves week to week, and how to separate drift from noise in a time series, is a different problem with a different answer.
What the best available data actually shows
The most careful public study of this is Don’t Measure Once: Measuring Visibility in AI Search (GEO), published by Julius Schulte, Malte Bleeker and Philipp Kaufmann at the University of St. Gallen on 10 April 2026 (arXiv 2604.07585). Its design is worth stating because the design is what makes the finding usable.
The authors took four Swiss-German verticals — telecoms, real estate, sporting goods, consumer electronics — built eight prompts for each, and ran them daily across ChatGPT, Gemini, Google AI Mode and Perplexity for 45 days between 24 January and 20 March 2026. Then a second collection: the same prompts fired up to ten times in immediate succession, so disagreement between runs could not be blamed on the index changing underneath.
The headline result is the one that got quoted. Day-to-day overlap in cited sources, measured with Jaccard similarity — the share of items two sets hold in common — averaged 0.336 to 0.423. Roughly 65% of the sources cited on Monday were not cited on Tuesday.
The result that did not get quoted is the important one. When the authors compared runs fired minutes apart on the same day, source overlap came back at 0.321 to 0.434 — the same range as the day-to-day figures. Runs two minutes apart disagreed about as much as runs twenty-four hours apart. Whatever moves between Monday and Tuesday is mostly what moves between 10:04 and 10:06, so index freshness and competitor activity are not the story these numbers have been used to tell.
What is prompt-level heterogeneity?
Prompt-level heterogeneity is the finding that different prompts carry wildly different amounts of instability — same category, same engine, same day. The variance in your report is therefore not one property of the engine that can be attached to your score as an error bar.
The St. Gallen team measured it directly: within a single campaign, per-prompt similarity ranged from above 0.8 to below 0.2. Some prompts returned near-identical brand sets on every repetition; others changed almost entirely. That is a range from essentially deterministic to essentially a lottery inside one eight-prompt basket.
There are two findings here and the industry built its playbook on the first. One: any single reading is unreliable. Two: the unreliability is distributed unevenly across your prompts. The first motivates more sampling. The second says the object you are sampling is not one thing, which is far less comfortable, because no amount of sampling fixes it.
What precision actually costs
It is worth pricing the sampling remedy before rejecting it. At a single run, the standard error of an estimated per-brand detection rate is 0.370 — a nominal 95% interval of plus or minus 72 points, which the authors note has to be clipped to the 0-to-1 range before it means anything. At seven runs it falls to 0.081, an interval of about plus or minus 16 points, which is where the paper’s recommendation lands: seven runs per prompt per day, eight where you care which URLs were cited.
Apply that to a real basket. A 150-prompt estate across four engines at seven runs is 4,200 calls a day, roughly 1.5 million a year — and on any individual prompt you still hold an interval of plus or minus 16 points. Reaching plus or minus 5 points on one prompt takes 384 runs of it on that engine. Per-prompt precision is not something agency budgets can buy.
Key takeaway. A single reading has an interval so wide it is arithmetically meaningless, and the run counts that fix it are unaffordable at prompt level. That is the first reason to stop chasing precision on the mean and start reading the shape instead.
Inclusion is binary, so the spread is not independent information
Here is the structural fact the measurement conversation keeps stepping over. In classical search a page holds a position on a scale and can move from 4 to 7 and back. In a generated answer there is no scale: as Wen and colleagues put it in their 2026 position paper, the competition is over inclusion and prominence inside the answer rather than ranked link positions. A source is in or out.
That makes each run of each prompt a Bernoulli trial — one yes-or-no draw with probability p. And a Bernoulli trial has a property fatal to the way visibility variance is reported: its variance is p(1 − p). The spread is a fixed function of the mean.
So reporting “our visibility is 40%, with a standard deviation across runs of 0.49” is reporting 0.4 and then reporting 0.4 again in a different hat. At prompt level the variation contains no information the mean does not already hold, because it was never free to be anything else.
Two consequences follow, and both have been circulating as though they were discoveries about how engines behave.
- Absolute swing peaks at 50%. In percentage points, which is how every dashboard reports it, the noisiest brand in a category is the one in the middle of the range. A brand improving from 10% to 50% shows up as having become less stable; nothing about the engine changed, p(1 − p) went from 0.09 to 0.25. The complaint “our score is so volatile” is loudest at exactly the level everyone is working towards.
- Relative swing is worst at the bottom. At a true rate of 4% the standard deviation of a single draw is nearly five times the mean, and separating 4% from 2% takes a sample most teams will never run. The brands most likely to be told their number jumped have the least number to jump.
Key takeaway. Any variance you report at the level of a single prompt is arithmetic, not evidence. If your provider presents within-prompt spread as a finding about engine stability, they are presenting p(1 − p) as insight.
The Dispersion Check: the only informative spread is between prompts
If within-prompt spread is uninformative, where is the information? Between prompts. There is a cheap test for whether your basket behaves like one coin or a mixture of several — one spreadsheet column, run once — and the answer changes what your report should say.
If every prompt shared a single underlying probability, the variance of your observed per-prompt hit rates would be predictable in advance: p̄(1 − p̄) divided by your run count k. Anything above that is overdispersion — spread a single coin cannot produce, and therefore evidence that your prompts are not draws from one population.
THE DISPERSION CHECK
1. Run each of your N prompts k times on one engine. Record each prompt’s hit rate: hits ÷ k.
2. Compute p̄, the mean of those N rates, and the sample variance of the N rates.
3. φ = variance ÷ [ p̄(1 − p̄) ÷ k ]. Published as a reproducible calculator, this is also the kind of asset that earns links on its own.
Reading it. φ near 1: one coin, and the mean is a genuine per-prompt probability you may reason with. φ between 1.5 and 3: mixed, and the mean is a poor summary of any single prompt. φ above 3: your basket is not one population, and every decision built on the mean is being made about an average of incompatible things.
A demonstration on public figures
AirOps reported two figures: roughly 30% of brands were visible from one answer to the next, and roughly 20% across five consecutive runs. That is enough to run the argument. The arithmetic below is mine, not theirs.
If one probability governed every brand, two-run consistency of 0.30 implies p = 0.548, and five-run consistency would be 0.548 to the fifth power: 4.9%. The observed figure is 20%, four times the homogeneous prediction — far too high for a population of identical coins. Fit the simplest mixture instead, a share of brands at effectively certain inclusion and the rest sharing one probability, and it lands near a fifth locked in with the remaining four-fifths on roughly a one-in-three chance.
The honest caveat: two summary statistics cannot identify a mixture, and I am fitting two parameters to two numbers, so treat the split as illustrative. The direction is not. Under any homogeneous model consistency decays geometrically with runs. It does not. The only structure that holds a five-run figure up at 20% is a locked mass — and that mass is invisible in the headline percentage.
It also explains an artefact that has caused a lot of panic: “consistency” cannot be reported without stating how many runs produced it, because it falls by construction as repetitions rise. Any brand can be made to look unreliable by a vendor with a larger sample. Before accepting a consistency figure, ask for k.
The Visibility Shape: a mean of 34% usually means nothing is at 34%
Once you accept that the basket is a mixture, the reporting problem resolves itself: stop summarising the distribution and partition it instead. Any real estate produces five recognisable bands, and they are not points on one scale — they are different situations with different causes and remedies.
| Band (hits over k runs) | What it actually is | What moves it | Can your own site move it? |
| Locked — k of k | In the retrieved pool and winning the slot on every draw | Nothing you can buy; only erosion by a stronger substitute | No — at ceiling |
| Dominant — 80%+ | Winning most draws, losing sometimes to a substitute | Reducing how substitutable you are for the runner-up | Partly |
| Contested — 20–80% | Retrieved, then selected some of the time. A selection problem | Marginal corroboration; being the non-redundant option | Partly — page work helps once you are in the pool |
| Fringe — under 20% | Occasionally retrieved, rarely chosen | Whatever raises retrieval frequency for that prompt | Rarely |
| Absent — 0 of k | Never in the retrieved set at all. An eligibility problem | Presence in the documents retrieved for that prompt | No |
What matters is not that the bands exist but that the three masses have different owners. The locked band is owned by you and by history, and marginal spend there returns nothing — unfortunate, since it is the band every dashboard screenshot is taken from. The contested band is where selection happens and the cheapest place to move a mean. The absent band is the only one where a document on somebody else’s domain is the sole available instrument, and it contributes exactly zero to the mean, which is why it is invisible in the report that sets the budget.
What is a structural zero?
A structural zero is a prompt where you were never in the retrieved candidate set, so no draw could ever have gone your way. It is categorically different from a low score, where you were retrieved and lost the selection.
On a dashboard, 0% and 5% differ by five points. In the pipeline they differ by a whole causal step. A 5% prompt says the retrieval layer surfaced you and the selection layer preferred something else — a problem that responds to being less substitutable, which is the mechanism behind how source diversity objectives pick between candidates. A 0% prompt says retrieval never surfaced you, and no page-level work reaches backwards to change whether that page was fetched — which is why on-domain remedies like internal PageRank sculpting cannot touch this band. Locating which layer failed is its own discipline, covered in the SERP-less audit method. The point here is narrower: a zero is not a small number, it is a different kind of object, and averaging it in with numbers is a category error.
The mean is a fine report and a bad instruction
The strongest objection to everything above is not that the shape is imaginary. The data says it is there. The strongest objection is that the shape does not matter to the decision, and it goes like this.
The objection. Our payoff is additive. We are cited or not, many times, across many prompts and many buyers. Expected appearances is just the mean times the number of draws, and expectation is linear — indifferent to the shape of whatever is being averaged. For the only question the board asks, which is how often a buyer sees us, the mean is not merely adequate but mathematically correct, and five bands are decoration.
That is correct as stated, and it deserves to be conceded rather than reframed. Where a payoff is genuinely additive in appearances, expectation is the right functional and the shape is irrelevant to it. Three things bound the concession.
- Thresholds break linearity. Most commercial consequences in this channel are lumpy rather than additive: making a three-name shortlist, being the recommendation rather than a mention, being the source that a research report carries into a document a buying committee reads. Where the payoff has a step in it, expectation is the wrong functional, and which side of the step you land on is decided by the shape. The compounding effect of being carried into downstream artefacts is the argument in how deep research modes turn citations into persistent documents.
- The weights are usually wrong. An unweighted basket assigns 1/N to every prompt, including prompts nobody asks, so the expectation is over your list rather than over demand. Fixing that is a sampling-design problem, and it relocates the error rather than removing it.
- The derivative differs by band. This is the real bound. Even where the mean is the right report, it is the wrong allocation signal, because the return on a pound is not the same across the three masses. In the worked example below, pushing the contested band adds 5.7 points and converting nine structural zeros adds 2.8, so a manager optimising the mean does the first every time. But contested gains are re-drawn on every query and decay as rivals move; a converted zero changes the support of the distribution and stays converted until the underlying document does. Equal expected gains, unequal half-lives, priced identically.
So the resolution is not “never report a mean”. It is narrower: report the mean, do not allocate on it. Put the five band counts beside it and both failure modes become visible to whoever holds the budget.
The second objection: determinism is coming
The other serious counter is that this is transitional: cached answers, seeded decoding and stabilised routing will squeeze the randomness out, and a score will be a score again. Three reasons that does not rescue the mean. Between-prompt heterogeneity is not a decoder property — it is a statement about which prompts retrieve you, and a deterministic engine still retrieves different sets for different questions, so determinism collapses the within-prompt variance that was already uninformative and leaves the informative part untouched. Caching hardens the shape rather than dissolving it, freezing a structural zero in place for the life of the cache. And personalisation runs the other way: a fully deterministic model with per-user context still produces a distribution across the user population, and your tool samples exactly one context — logged out, default settings, no memory, one location.
The Decision Functional: which summary your question actually needs
If visibility is a distribution, “what is our score” is an incomplete question in the way “what is the weather” is. A distribution has many summaries and the right one is chosen by the decision, not by the statistics. Five questions come up in practice; one is answered by the number on the dashboard.
THE DECISION FUNCTIONAL
“Are we in the game at all?” → the support: the share of prompts where p exceeds zero. A brand at 40% across 60 prompts and one at 40% across 140 are in different businesses.
“Will a given buyer see us?” → the demand-weighted mean — weighted by how often each prompt is asked, not by how many you wrote down.
“Can we promise a client we will appear?” → a lower quantile. A promise is about the bad draw, not the average one; an agency contracting on a mean is selling a number no client experiences.
“Is the position defensible?” → the mass at or near certainty, and its trend. This is the only band that survives a competitor’s campaign.
“Where is the next £10,000 productive?” → the contested and absent masses, priced separately: they buy different durations of the same headline movement.
Five questions, five statistics. The practical test of a reporting system is whether it can answer any of them without a rebuild, and most cannot, because they were built to produce a single index. That is the trap in most AI-era measurement tooling, and the reason two providers can hand the same client scores 32 points apart without either being wrong.
What Ofsted worked out about single numbers
England has just run this experiment at national scale on a different subject, and the sequence ends somewhere useful. For years, state schools received a single headline grade — Outstanding, Good, Requires Improvement or Inadequate — summarising a set of separate underlying judgements. The grade became load-bearing far beyond what it was designed to carry: house prices, admissions behaviour, recruitment, intervention decisions. Ofsted’s own research found that fewer than four in ten parents, and only 29% of teachers, supported one-word judgements.
The Department for Education removed single headline grades in September 2024, and Ofsted began inspecting under a report card from 10 November 2025, grading multiple evaluation areas on a five-point scale. The mechanism being corrected is precisely the one in this article: a heterogeneous set of judgements collapsed into a scalar for management convenience, with consequences attached to the scalar rather than the judgements underneath it.
Note what the fix was not. It was not a better single grade, and it was not a confidence interval around the old one. It was to publish the profile and let the composite be assembled downstream.
The honest counterpoint is the same objection as before, and it lands: readers immediately reconstruct a composite, and parents and the press were building their own rankings out of the areas within weeks. But that is not a failure of the reform. It is the difference between a collapse that happens in the open, using weights the reader chose and can argue with, and one that happens inside the instrument using weights nobody published. Hold your AI visibility report to that standard. If the board wants one number, give them one number — and the five band counts it came from, so the weighting is theirs, visible and contestable.
What the shape means for a link building budget
Three consequences follow for anyone allocating spend, and they are why this is a link building article rather than a statistics one.
Set the budget split by the concentration of the surface
The St. Gallen team also computed a Gini coefficient — a 0-to-1 measure of how unequally citations are shared out across domains — for every campaign and engine pairing. The mean was 0.715. Google AI Mode was the most concentrated at 0.782; Perplexity the least at 0.671.
Concentration describes everyone’s distribution at once. The more concentrated a surface, the more of the field sits in the absent band and the narrower the contested middle. On a high-concentration surface eligibility binds and selection work is a game only the already-cited play; on a low-concentration surface the reverse. The same budget buys different things on Google AI Mode than on Perplexity, and the concentration figure says which. Machine-readable estate such as AI-readable API feeds shifts retrieval frequency rather than selection, so it belongs to the eligibility budget too. Which surfaces reward which assets is tracked in the 2026 link building statistics and in the mechanics of how AI Overviews consume backlinks.
Enumerate the zeros, then name what was cited instead
For every prompt in the absent band, record which documents were retrieved. Not which competitors — which documents. That is a prospect list, produced by an audit you already paid for, and the only one in this discipline derived from a question a buyer asked rather than an index of who links to whom. It holds things a competitor backlink analysis never surfaces: trade registers, regulator pages, a specialist directory. Those documents are the eligibility instrument, and there is no on-site substitute, which is why what a backlink actually is still matters in a channel with no rankings, and why the scarcity of corroborated claims prices these placements as it does.
Prioritise inside that list by whether a placement is fetchable and durable rather than by domain metrics — and screen it the way you would screen for low-quality link patterns, because a retrieved page that gets discounted is worse than no page. A niche edit into a page already retrieved for the prompt converts a zero more reliably than a stronger link on a page that is not. That is a different screen from anything in the standard strategy set, and it can only be built from an audit that reports its zeros.
Price the two moves by half-life, not by points
Contested gains and converted zeros are not the same asset with different price tags. A contested gain is re-drawn on every query and decays as rivals publish, which is why news-pegged placements flatter a headline mean and rarely hold it; a converted zero changes the support of the distribution and persists until the underlying document changes or is removed. Any allocation model that prices them by their effect on the headline mean will systematically overbuy the perishable one, in the same way that session-based traffic models systematically defund the zero-referral sources doing the most work.
Two notes. A converted zero is only as durable as its host page, so the maintenance question is whether that page stays live, crawled and accurate — ordinary technical SEO hygiene. And converting zeros is slow, lumpy work that makes your link velocity look worse before it makes your shape look better: a month-one conversation, not a month-four one.
Key takeaway. The mean rewards you for work in the band where returns are lowest and durability is shortest, and gives you nothing for the band where an earned document is the only available instrument. That is not a reporting inconvenience. It is a budget being steered by a summary statistic that cannot see the thing it should be buying.
Worked example: Holbeck & Vane, Leeds
Holbeck & Vane is a commercial insurance broker in Leeds — 61 staff, £4.6M in brokerage income, strongest in fleet and property but making its margin on professional indemnity and cyber. The figures are constructed to be reproducible rather than reported; the arithmetic is the point.
February 2026. A visibility audit: 140 prompts, four engines, seven runs each. The headline came back at 34.6% against a competitor set, and the board read it as “we appear in about a third of AI answers about commercial broking”.
March 2026. The head of marketing ran the Dispersion Check on the same raw data in forty minutes. The bands: 19 locked at 7 of 7, 13 dominant at 6 of 7, 28 contested at 4 of 7, 16 fringe at 1 of 7, 64 absent at 0 of 7. Observed variance of the per-prompt rates was 0.153 against a homogeneous expectation of 0.032, so φ = 4.7. Nothing in the basket was at 34.6%; 46% of it was at zero.
March to May 2026. The retainer went where retainers go — into the contested band, because that is where the mean moves fastest. It worked: pushing 28 contested prompts from roughly 4 of 7 to 6 of 7 lifted the headline from 34.6% to 40.3%, reported in May as a win. Meanwhile the 22 prompts tied to professional indemnity and cyber sat at a weighted 10.4%, because 18 of the 22 were structural zeros and no contested work touched them.
June 2026. The team enumerated the 64 zeros and logged what had been cited instead. The list was not what a backlink tool would produce: a regional brokers’ register, two trade titles reached through expert-source platforms, an insurer’s technical guidance page, one FCA-adjacent explainer. Nine placements were commissioned against it over ten weeks, seven on prompts inside the professional-indemnity and cyber set.
August 2026. Nine zeros converted to roughly 3 of 7, lifting the headline mean by 2.8 points — half the movement the contested push produced, at comparable cost. On the headline it was the worse quarter. On the 22 commercially weighted prompts the figure moved from 10.4% to 24.0%. The spring’s contested gains had partly decayed as two rivals published; none of the nine converted zeros reverted.
Four honest negatives, because a worked example with no losses is a brochure. Two of the nine placements produced no measurable change — live and retrieved but never selected, a contested outcome dressed as an eligibility win. The reallocation cost the marketing lead a difficult July board meeting, with a flat headline and a distributional explanation. The φ statistic is engine-specific and the team first computed it pooled across all four, understating dispersion. And the whole exercise rests on seven runs, leaving each prompt an interval of roughly ± 16 points — fine for sorting a prompt into a band, useless for arguing that any single prompt moved.
The Monday checklist
Seven things, none of which require new tooling or a new vendor.
- Ask your provider for k. How many runs produced the number in your report, on each engine? If the answer is one or two, the interval around that number is wider than the number itself and nothing built on it is safe.
- Run the Dispersion Check on data you already have. Per-prompt hit rates, one variance, one division. Above φ = 3, stop reasoning with the mean this week.
- Partition the basket into the five bands and report the counts, not the average. Do it once by hand before asking anyone to build it.
- Count your structural zeros as a share of the basket. That figure describes your position better than the headline score, and predicts how much of the budget has to go off-domain.
- List what was cited instead, for every zero. Documents, not domains. This is your prospecting list for the quarter and it is already paid for.
- Split the budget explicitly between selection work in the contested band and eligibility work on the zeros, writing down the expected half-life of each before you spend so nobody re-litigates it in month three.
- Rewrite the KPI so it names a band count alongside any mean. “Reduce structural zeros from 64 to 45” is actionable. “Raise visibility from 34% to 40%” will be met in the cheapest and least durable way available.
The sentence the field settled on is correct: visibility is a distribution, not a number. The mistake was hearing it as a warning about precision. It is a statement about kind. Take it literally and the practical question stops being how confident you are in your score and becomes which parts of your market you are eligible to appear in at all — a question about other people’s documents, and always was.
