TL;DR
Your score moves for four reasons and only one of them is you. On the 2026 evidence, the three you do not control dominate at any horizon shorter than a quarter.
The cleanest test in the literature is the simplest. Swiss researchers re-ran the same prompts twice in the same hour: two simultaneous runs disagreed as much as runs 45 days apart. A gap that does not grow with time is not drift.
The citation half-life genre applies survival analysis to a process with no absorbing state. Around 57% of brands that vanish from one answer come back in a later one. You cannot fit a decay rate to something that resurrects.
Detection is paid for in the one currency you cannot buy more of. On the best published standard errors, two 14-day windows detect a 22-point change and nothing smaller; a five-point gain needs roughly six months.
That is longer than your intervention cycle, which makes monthly reallocation Deming’s funnel Rule 2: adjusting a stable process on its own noise, which doubles the variance of the output. Some of the instability the field complains about is manufactured by the field.
Steer on the input. The corroboration record is the only quantity here observed on the day it happens, with no sampling error, no revision and no vendor basket under it.
The number moved. That is not information.
The score was 24 last month. It is 19 today. Somebody asks what happened, and the honest answer, which nobody gives, is that the question is out of order. Before why it moved comes whether it moved, and almost no team in 2026 has an instrument capable of answering the second question, let alone the first.
Three words get used interchangeably in that meeting and they are three different objects. Variance, the spread you would see across repeated observations of the same thing with nothing changed. Drift, a slow directional change in the underlying system. Noise, movement that carries no information about the thing you are trying to steer. The standing advice in the field is to check weekly and wait for a two-week trend before acting, which assumes you are looking at drift. The evidence says you are almost always looking at variance, and that the two-week trend is a fiction with a decimal point attached.
What is AI visibility variance? It is the spread in your measured visibility across repeated observations of the same brand, prompts and engines. It has four sources: sampling noise, composition churn, structural breaks in the platform, and your own work. Published 2026 estimates put the last well below the first three over any period a marketing team would call a reporting cycle.
The test that separates drift from noise
A drifting process leaves a fingerprint: dissimilarity between two observations grows with the gap between them. In a random walk the variance of the difference is proportional to elapsed time, so two readings an hour apart should agree more closely than two a month apart. Geostatisticians formalised this as the variogram, a plot of dissimilarity against lag. A rising line means the system is moving. A line flat from lag zero means the disagreement was always there and time has nothing to do with it.
Nobody in generative engine optimisation seems to have drawn it, which is unfortunate, because the data arrived in April 2026. Schulte, Bleeker and Kaufmann at the University of St Gallen ran four Swiss-German verticals, eight prompts each, across ChatGPT, Gemini, Google AI Mode and Perplexity, daily for 45 days, plus up to ten same-day repeats of each prompt. Day-to-day source overlap, as Jaccard similarity (the share of sources two answers hold in common), ran 0.336 to 0.423. Same-day simultaneous repeats ran 0.321 to 0.434.
Those are the same range. Two answers generated seconds apart disagree about their sources as much as two generated six weeks apart. Whatever is moving your weekly chart did not accumulate over the week. The paper’s headline was that you should not measure once; the buried result is the more useful one, and the field has not absorbed it: at these horizons, the lag does not matter.
What a real drift term looks like
This is not a claim that nothing moves. Real, dated, directional shifts exist and they are large. Semrush’s 13-week tracking caught ChatGPT’s Reddit citation share falling from roughly 60% to roughly 10% within weeks in September 2025. Growth Memo put ChatGPT’s share of AI-search activity at 56% in its H1 2026 report, down from 78% six months earlier. Semrush’s clickstream analysis has the share of ChatGPT queries that trigger a live web search falling from 46% to 34.5% between October 2025 and February 2026, which mechanically reduces how often anybody gets cited and quietly rewrites the zero-click traffic model underneath every forecast.
Every one is a genuine change in the system and not one is about your content. Each also carries the same plain signature: it moves the competitors in your set at the same time and in the same direction.
Key takeaway
If a move is common to your competitors, it is not about your content, whichever direction it went. Check two rivals before you check your own pages.
The half-life industry measured the machine, not your content
The dominant 2026 explanation for a falling score is citation decay, and it arrived with impressive numbers: a median citation half-life around 4.5 weeks across 3.5 million events at Scrunch and Stacker and again across 240 million at Profound, 40 to 60% of cited domains rotating month to month, and a Citation Retention Rate averaging 33% over 28 days at Digital Authority Partners, which put AI Overviews at 27% and Gemini lowest at 11%. Table 1 sets out what each study measured.
The sentence that dismantles the genre
Every one of those studies is competently run. The inference drawn from them, that your pages are decaying and need a refresh treadmill, follows from none of them, and the study that shows why is one of the four. Digital Authority Partners ran three waves across six weeks. Two-thirds of the cited sources changed. The rate itself, the share of answers in which a brand appeared, held inside a three-point band with no visible drift in any direction. Membership churned violently while the level did not move at all.
That is not decay. It is turnover in a stationary process, and the two are identical to anyone tracking named sources without tracking the level. The half-life framing quietly imports a property a citation does not have: that you keep it until something takes it.
Survival analysis needs an absorbing state
A half-life is a parameter of a decay process, and a decay process needs an absorbing state: once you leave, you are gone. Citations have no such state. AirOps’ analysis of more than 45,000 citations found that around 57% of brands that disappeared from one response reappeared in a subsequent one. Death is the wrong model for something that comes back better than half the time.
Run the numbers on the published figures
What a survival curve fits, in a process with no absorbing state, is the reciprocal of a re-draw probability. If an answer fills k citation slots from an eligible pool of N broadly interchangeable sources, the chance any given source returns is about k divided by N. Perplexity showed 5.9 citations per query and the highest retention of the five engines at 44%: the implied pool is about 13 domains. Run it backwards for a three-slot answer at the 33% cross-engine average and the pool is about nine. The engine showing the most sources per answer holds the most between answers, which is what a slot-count model predicts, what the way an answer assembles its source set implies, and what a decay model has no reason to expect.
Now compare the observed half-lives against the right null. The field’s implicit null is permanence: 100% retention, the way a ranking behaves. The correct null is re-draw. At 33% retention per 28-day wave, continuous presence halves in about 0.63 of a wave, roughly 18 days, against published cohort half-lives of 3.4 to 5.8 weeks. Measured persistence sits between one and two times the pure-chance baseline and nowhere near permanence, so most of the measured decay is the distance between the wrong null and the right one.
There is a second tell. Estimates of the same quantity differ threefold across studies, from MaxAEO’s ten days to Scrunch’s 4.5 weeks, and the gap tracks observation cadence rather than vertical or content type. A statistic that shifts when you change how often you look is measuring the shape of the answer, not the quality of your page.
Table 1. The 2026 persistence estimates and what each one actually measured.
| Study and scale | What was tracked | Reported figure | What it is evidence of |
| Scrunch / Stacker, 3.5m citation events, Sep 25 to Mar 26 | Cohort survival of cited sources across weeks | Median half-life 4.5 weeks; ChatGPT 3.4, Perplexity 5.7 | Persistence above the re-draw baseline, below permanence |
| Profound, 240m citations | Month-to-month rotation of cited domains | 40 to 60% rotate monthly, 70 to 90% over six months | Composition churn; says nothing about the level |
| MaxAEO, 45-day durability run | Daily observation of first-citation cohorts | 58% alive at day 7, 39% at day 14, half-life near day 10 | Cadence effect: look more often, measure a shorter life |
| Digital Authority Partners, 3 waves, 5 engines | Retention of cited URLs and the underlying rate | 33% retention over 28 days; rate flat inside a 3-point band | The decisive result: churn without decline |
| AirOps, 45,000+ citations | Reappearance after disappearance | ~57% of vanished brands return in a later response | No absorbing state, so no valid half-life |
Instrument one: the re-draw baseline
THE RE-DRAW BASELINE
1. Take one prompt on one engine and run it ten times inside a single hour. Record every domain cited in every answer.
2. Set k as the mean number of citation slots per answer. Set N as the count of distinct domains that appeared at least once across the ten runs.
3. Baseline single-wave retention is k divided by N. Baseline continuous-presence half-life is ln(0.5) divided by ln(k/N), expressed in waves.
4. Compare that against the retention figure your dashboard reports over the same interval.
5. Read the verdict. Observed within a few points of k/N means the movement is re-draw: there is nothing to explain and nothing to fix. Materially above k/N means you hold an anchored slot, and the job is to protect the documents that put you there. Materially below k/N is the only case warranting a hunt for a cause, and it begins with a date rather than a page: a platform change, a licence expiry, or a manual action.
Limit. No vendor publishes N, because N is a property of the pool rather than of you. You have to run raw prompts yourself, with or without the tools in your stack, which is why it is worth more than anything on the dashboard.
Four causes, four owners, and the one that is yours
Sampling noise
Test: re-run the identical prompt within the hour. Owner: nobody. Permitted action: none, ever. This is the largest single component in every published dataset and the one most often written up as a content problem. When a tracker shows a brand vanishing between two checks an hour apart, audit the method before you audit the marketing.
Composition churn
Test: the level holds while the named sources rotate. Owner: the slot architecture of the answer, set by the engine and different on every surface, so deep research modes with dozens of slots churn far less than a three-slot chat reply. Permitted action: increase the number of independent documents capable of filling your slot, which raises the probability that at least one of them is drawn. Not: replace the page that dropped out.
Structural break
Test: dated, discontinuous, and common to your competitors. Owner: the platform, or your own vendor. A basket expansion or a re-baselining is indistinguishable in the chart from a category collapse, so demand two dates from any provider: the last basket change and the last re-baselining. A series that crosses either is two series wearing one line.
Your own effect
Test: it survives the detection calendar below. Permitted action: this is the only movement that licenses reallocating budget, and in most months it does not survive. The correct report is a sentence saying so rather than a narrative fitted to the wiggle, and a presence audit that localises absence is a better use of the hour than an explanation of a five-point dip.
The cost of certainty is time
The St Gallen paper includes a table almost nobody has quoted: rolling-window standard errors for a brand detection rate as the window widens. At one day, 0.322; at seven days 0.135; at ten 0.107; at 14 0.080; at 21 0.053; at 28 days 0.033. Those are the honest error bars on the number in your deck, and on most of the benchmark figures quoted beside it.
Turn them into something a budget holder can use. Claiming a change means comparing two windows, so the relevant quantity is the standard error of the difference, the square root of two times a single window’s, with a 95% threshold at 1.96 times that. Two 7-day windows detect a 37-point change and nothing smaller. Two 14-day windows detect 22 points. Two 21-day windows detect 15. Two 28-day windows detect 9.
So the advice to expect five to ten per cent of weekly noise and wait for a two-week trend before changing strategy, which circulates in almost every 2026 measurement guide, understates the requirement by a factor of two to four. At 14 days the smallest readable change is roughly 22 points, and a brand moving 22 points in a fortnight has not been optimised. Something happened to the platform.
Why halving the effect quadruples the wait
Standard error falls with the square root of the window, so the window required to see a given change rises with the inverse square of that change. Halve the effect you want to detect and you quadruple the time it takes. Extrapolating the published series on that basis, which is my arithmetic rather than the authors’, a five-point improvement needs windows of about 13 weeks on each side, roughly six months elapsed. A three-point improvement needs about 37 weeks a side, or the better part of 17 months.
How long before I can say my score moved? Read it off the calendar below. If the change you care about is under nine points, the answer is longer than a quarter, and reporting more often will not shorten it.
Table 2. THE DETECTION CALENDAR. Windows required at a 95% threshold, derived from the St Gallen rolling-window standard errors. The last column assumes one substantive change a fortnight.
| Change you want to detect | Window each side | Elapsed before you can speak | Changes shipped meanwhile |
| 37 points | 7 days | 2 weeks | 1 |
| 22 points | 14 days | 4 weeks | 2 |
| 15 points | 21 days | 6 weeks | 3 |
| 9 points | 28 days | 8 weeks | 4 |
| 5 points | 13 weeks | 6 months | 13 |
| 3 points | 37 weeks | 17 months | 37 |
Limit worth stating. The published errors come from a table the authors flag as flattered by a finite-population correction at high run counts, the study ran on Swiss servers in German, and the extrapolation past 28 days is mine. Treat the bottom two rows as optimistic, which makes the argument stronger rather than weaker.
Dead time: the loop nobody drew
Put the calendar next to your operating rhythm and the real problem appears. Three lags separate an action from its observable consequence: production, from commissioning work to it existing in public; propagation, from existing to being in the pool the engine draws from; and detection, the windows above. The field debates the first, occasionally worries about the second, which is what technical work on crawl and discovery actually shortens, and almost never counts the third, which is the largest.
Control engineers have a name for the total: dead time, the delay between acting and being able to observe the result. The relevant result is unkind. When dead time is long relative to the system’s own response time, no tuning of the feedback loop makes it stable, and tightening the loop by correcting faster makes the oscillation worse rather than better. The remedy is not a better controller. It is feedforward: steering on a measured input rather than a measured output.
A monthly review cycle acting on a variable with an eight-week detection window is a loop actuated twice as fast as it can be read. It will oscillate. It is not badly run and the people in it are not careless; it is badly specified, and no amount of diligence inside a misspecified loop produces convergence.
Tampering, or how a field manufactures its own variance
W. Edwards Deming demonstrated this with a funnel, a marble and a target. Rule 1 is to leave the funnel alone. Rule 2 is to move it from its last position by the size of the last error, which is what every sensible person does. Rule 2 keeps the process stable and doubles the variance of the result, spreading the scatter about 41% wider than doing nothing; Rules 3 and 4 diverge without limit. His summary in Out of the Crisis is blunt: adjust a stable process because a result was undesirable, or unusually good, and what follows will be worse than leaving it alone.
Monthly reallocation against a visibility score is Rule 2 with a slide deck. Each decision is defensible alone and the sequence has a measurable cost: it converts a stable programme into a random walk, and it guarantees no tactic is ever exposed long enough to be evaluated by the calendar above.
The NHS ran this experiment first
NHS Improvement published Making Data Count in May 2018 for exactly this failure, naming two-point comparisons and red-amber-green reports as the practices to abandon and statistical process control as the replacement. A 2020 review of board papers across every hospital trust in England still found 75 relying predominantly on RAG ratings or two-point comparisons.
Then the coda, which is the most honest thing here. A parallel cluster randomised trial across 15 English hospitals tested whether the training worked. Control chart use rose 29 percentage points in the trained arm and 30 in the untrained one, an absolute difference of 9% with a 95% confidence interval running from minus 34% to plus 52%. The programme designed to stop people over-reading noise could not itself be distinguished from noise. Hold that interval in mind the next time a dashboard reports a five-point move without one.
THE TAMPERING AUDIT
1. Count the allocation changes in the last 12 months that followed a movement smaller than your detection threshold. Every one of them is Rule 2.
2. Compute the mean uninterrupted exposure of a tactic before it was changed. If it is shorter than the window in the calendar, nothing in your programme has ever been evaluated, including the things that worked.
3. Check whether any reading in your history has an interval attached. If none does, none of them are measurements, and the series is a sequence of anecdotes in a line chart.
4. List everything that varies between runs: prompt wording, session state, location, login, engine version. MaxAEO could not reproduce the widely quoted 30% answer-to-answer figure under controlled conditions and puts the gap down to exactly this. Variance manufactured by the instrument is charged to your content in the write-up.
Verdict. Zero sub-threshold changes is disciplined. One to three is ordinary. Four or more and you are the largest single source of variance in your own programme, ahead of the engines you are complaining about.
The objection that survives
The strongest counter is not that the maths is wrong. It is this: we know the score is noisy, we ignore small moves, we act on big ones, and big ones are real. That is what the better practitioners do, and the arithmetic above supports rather than refutes it. A 25-point drop clears the fortnightly threshold comfortably; it happened. It survives, bounded three ways.
First, the big moves are disproportionately not yours. Reddit’s share inside ChatGPT falling from 60% to 10% in weeks; a 22-point swing in one engine’s share of all AI-search activity inside a single half-year; a vendor expanding its prompt basket. Platform-level events move every brand in the category simultaneously, which is exactly what makes them the largest moves in any series. A policy of acting only on large movements is, by construction, a policy of acting mostly on other people’s events.
Second, the action a big drop triggers is asymmetric. A large fall prompts stopping something rather than starting something, and stopping earned-placement work is not the reverse of starting it: the stock already built keeps working while the cancelled flow takes a quarter to restart. A false positive on a drop costs more than a false negative on a gain.
Third, allocation compares two small numbers, not one big one. You are rarely choosing between doing something and doing nothing, but between two credible programmes whose expected difference is a few points by construction: precisely the region the instrument cannot resolve. A signal coarse enough to be safe for reporting is the wrong instrument for the decision it is used to make.
Key takeaway
Use the visibility score to detect events. Do not use it to grade work. Those are different jobs and only the first one is inside the instrument’s resolution.
What this changes for earned links
The corroboration record is the only undelayed variable in the system
If the output cannot be read faster than you act, you steer on an input, and which input follows from the properties the score lacks rather than from any claim about what a link is worth.
A third-party document either exists on 14 September or it does not. No sampling error, because you are counting rather than estimating. No revision, because last quarter’s count reads the same next year and the record is dated and non-repudiable. No vendor basket underneath it and no re-baselining that splits the series in two. Against a number whose honest error bar is plus or minus 26 points on a weekly window, the record of who has linked to you and when is the only quantity here that can be read on the day it changes.
The anchor arithmetic
Composition churn has a defence and it is not better writing. If each document that supports your claim has probability p of being pulled into a given answer, and you are supported by m broadly independent documents, your chance of appearing at all is 1 minus (1 minus p) to the power m. At p = 0.33, that is 33% with one document, 55% with two, 70% with three, 87% with five, 91% with six and 96% with nine.
Two things fall out. One document to three buys 37 points of appearance probability; six to nine buys five. And the anchored brand, the 2.3 per prompt that MaxAEO found holding 80% of daily snapshots while four to nine challengers rotate through the remaining slots, is not the one with the best page. It is the one the engine can support from whichever sources it happens to draw. Anchoring is a redundancy property, and redundancy is bought off your own domain: no quantity of internal link engineering substitutes for a second publisher.
Sizing rule. Five to six genuinely independent corroborating documents per claim cluster puts you in anchor territory. Below three you are running a coin flip. Above nine you are buying rounding error, and whoever sold you the tenth knows it. The load-bearing word is independent: twelve outlets carrying one agency’s syndicated copy, or one quote placed through a journalist-request platform and picked up widely, is one document with twelve URLs, and the arithmetic collapses accordingly.
Batch the earning, do not smooth it
Detection cost scales with the inverse square of the effect, which makes campaign sizing a measurement decision rather than a cash-flow one and reorders the tactic mix by batch size rather than by cost per link. A programme shipping two or three placements a month produces a series in which nothing is ever readable, however good the placements are. The same annual budget concentrated into two or three dated waves clears the threshold, and because the waves are dated you can put measurement windows either side of them instead of averaging across them.
That cuts against the standing advice to keep acquisition smooth, and the tension is real. The link velocity literature addresses a manipulation-detection problem and says little about earned, dated coverage arriving in clusters, which is how coverage has always arrived when something actually happened. Honest position: the evidence for velocity penalties on earned coverage is weak but not zero, so batch the earning and leave the anchor text alone.
A prospecting screen that follows from the maths
Prefer host documents that will still be in the pool when the engine re-draws tomorrow. A register entry, a standards body page, a machine-readable reference feed and a calculator or case-study asset on a stable URL persist by design; a slot in a rotating listicle is itself subject to re-draw, so you are buying a probability of a probability. The screen is cheap: ask whether the page has changed its own content in the last year, and prefer the ones that have not.
Worked example: Duddon Fire Engineering
Duddon Fire Engineering is a passive fire protection contractor in Barrow-in-Furness, 61 staff and £9.4m of turnover, selling to housing associations and main contractors whose technical teams now open a specification question in an assistant before they open a browser. It retained a generative-visibility vendor at £6,800 a month from September 2025.
The first four dashboard readings were 24, 19, 27 and 21. Each was reviewed and each triggered a change: schema work in November because November was down, a refresh sprint in December because December recovered, a community push in January because January dipped again. Every decision was reasonable and the sequence was Rule 2.
In March they ran the re-draw baseline themselves. Ten runs of one core prompt gave a mean of 3.4 citation slots and 11 distinct domains, a baseline retention of 31% against the 34% their dashboard reported. The calendar said four monthly readings could not resolve anything smaller than 22 points, and the largest gap in the series was eight. Nothing had happened in either direction, and £27,200 had been spent chasing the difference.
What changed: monthly reallocation stopped. One 14-week programme ran instead, with placements batched into two dated waves in March and June of nine and eleven independent documents, across a fire-safety trade title, two housing-association bulletins, a Building Safety Act commentary series, a manufacturer’s case-study library and two regional business titles, each carrying an independently checkable claim rather than a restatement of the others. Measurement moved to a 28-day window either side of each wave.
The result, stated honestly. The pre-wave 28-day reading was 22.6 and the post-wave reading 31.4, a gap of 8.8 against a threshold of 9.1: just under, and therefore not claimable. The second wave read 31.4 to 39.2, a gap of 7.8, also under threshold. Only the full-year comparison, 22.6 against 39.2, cleared it. The independent-document count went from 7 to 27, readable on the day it changed. The largest single jump in the series, nine points in one week in May, coincided with an engine changing its live-search default and moved two competitors similarly, so it was written up as an event rather than a win. The lesson they drew was that the waves were still too small: one wave of twenty would have cleared the threshold alone.
The Monday checklist
- Run one prompt ten times in one hour on one engine. Record slots and distinct domains, compute k/N, and put it at the top of the reporting template. Everything else is read against it.
- Put an interval on every number. If you cannot compute one, write no interval beside it. That is more honest than a decimal point.
- Set the reporting cadence from the calendar, not from the invoice cycle. If the smallest change worth acting on needs 28 days, stop reporting weekly.
- Count your Rule 2 decisions over the last 12 months and stop making them. Publish the count internally.
- Demand two dates from your vendor: the last prompt basket change and the last re-baselining. Split the series at both.
- Replace the score as your steering variable with a dated count of independent third-party documents corroborating each claim cluster you care about.
- Compute 1 minus (1 minus p) to the power m for your top three claim clusters. Any cluster sitting below three independent documents is where next quarter’s budget goes.
- Batch the next two quarters of earned placement into two dated waves with a 28-day window either side of each, then leave it alone until the window closes.
None of this makes the number more accurate. It makes it honest, a different and more useful property, and it moves the steering onto a variable that answers on the day you ask it.
