TL;DR
Every AI visibility audit carries two errors. One is sampling error, which more runs genuinely fix. The other is frame error, the gap between the prompts in your basket and the questions your buyers actually ask, and no run count fixes it, because more data makes a badly chosen basket more precisely wrong. Semrush’s own clickstream research found that 65 to 85 per cent of ChatGPT prompts had no match in a 27-billion-keyword database, and that the ones which did match skewed navigational and transactional. Every standard sourcing recipe therefore biases in the same direction: your score reads too high, and your instrument is blind to exactly the work most likely to move it. This is a design problem, and the fixes are structural.
Two errors, and only one of them is for sale
By mid-2026 the field had converged on an answer to audit design, and the answer was a number. MaxAEO recommends 40 to 100 buyer-intent prompts and 300 to 900 prompt-runs per engine per test period. A mention.network guide of 17 July 2026 recommends 20 to 30 buyer-language prompts across four intent groups, frozen wording, trend read after three runs. Neil Patel puts it at 15 to 30 and adds that representativeness matters more than volume. Click Laboratory recommends 20 to 40 for a mid-size B2B site, and is admirably honest about why: more than that and the monthly recheck stalls.
Every one of those four answers the question how many, and the ceiling in at least one case is set by operational load rather than by any statement about what the resulting number estimates. Representativeness gets invoked. Nothing in the genre defines the population it is meant to represent.
An audit has two error terms. The first is sampling error: you asked a finite number of questions a finite number of times, and a different draw would have given a slightly different answer. It shrinks roughly with the square root of the sample, and you can buy it down with API calls. The second is frame error: the distance between the questions you put in the basket and the questions your buyers put to an assistant. It does not shrink with sample size. It cannot be bought down at any price.
The 2026 methodology conversation is almost entirely about the first term. The money is in the second.
What is prompt-space sampling?
Prompt-space sampling is the practice of choosing a finite set of prompts to stand in for the unbounded set of questions real buyers put to AI assistants, and then stating what the resulting score is a measurement of. It involves three decisions: which prompts go in, what weight each carries, and how a fixed call budget divides across prompts, engines and repeat runs. Only the third is a sample-size question, and it is the least consequential of the three.
That ordering falls out of arithmetic long settled in survey statistics and not yet carried across into measuring authority in AI systems.
The arithmetic that makes the basket dominate the run count
In 2018 Xiao-Li Meng published an identity in the Annals of Applied Statistics that decomposes the error of any sample average, probabilistic or not, into exactly three factors. The first is a data quality term, the correlation between whether a unit is in your data and the value you are measuring on it. Meng calls it the data defect correlation, written rho. The second is a data quantity term. The third is problem difficulty, the spread of the thing being measured.
The consequence is uncomfortable. Probability sampling holds rho near zero by design. Once you lose that control, the error grows with the square root of the population rather than shrinking with the square root of the sample. Meng named the result the Law of Large Populations, and summarised its practical form in a sentence worth keeping on a wall: the more the data, the surer we fool ourselves.
INSTRUMENT 1 — THE EFFECTIVE BASKET SIZE
n_eff = n / [ 1 + (N – n) x rho squared ]
n is the number of prompts in your basket. N is the size of the prompt population you claim to describe. Rho is the correlation between being in your basket and being a prompt where you get cited.
Check it against Meng’s worked case. Around 2.3 million self-reported presidential preferences, roughly one per cent of the American electorate, with a defect correlation of 0.005, carry the same mean squared error as a random sample of about 400 people: a 99.98 per cent reduction produced by a correlation of half of one per cent.
Now a citation audit. Take a 300-prompt basket, a category in which perhaps 25,000 distinct questions could plausibly be asked, and a defect correlation of 0.05. The denominator is 1 + 24,700 x 0.0025, which is 62.75. The effective basket size is about five prompts.
The tolerable defect. Invert it. For that basket to behave like 100 random draws, rho must sit below about 0.009. Ask yourself honestly whether a basket built from your own keyword export is within one per cent of independent of where you already appear.
Bands: below 0.01, the basket is behaving like a sample. Between 0.01 and 0.05, the interval your vendor prints is fiction but the ordering of strata probably survives. Above 0.05, the level is not identified and only shape and direction can be reported.
The empirical demonstration everyone should read is Bradley and colleagues in Nature in 2021. Delphi-Facebook was collecting roughly 250,000 responses a week, over 4.5 million across nineteen waves. In May 2021 it overstated American first-dose vaccine uptake by 17 percentage points, with a bias-adjusted effective sample size of under ten. The Census Household Pulse survey, at about 75,000 responses a wave, was out by 14. An Axios-Ipsos panel of roughly 1,000 responses a week, run to established survey practice, got it right and published honest intervals.
The two surveys did not differ in effort or budget. They differed in design.
How many prompts should an AI visibility audit track?
Enough that every stratum you intend to report separately has a floor of roughly 15 to 20 prompts, and no more than you can afford to re-source on a schedule. The count is not the binding decision. A 400-prompt basket pulled from a single keyword export carries less information than a 120-prompt basket assembled from four independent sources, and it costs more to run.
Every recommended sourcing recipe biases in the same direction
SE Ranking’s April 2026 guide to choosing prompts lists the sourcing methods the field actually uses: convert your existing SEO keywords, mine People Also Ask and AI Overviews, analyse Reddit and community forums, reverse-engineer competitor websites, and start with the keywords your competitors are paying for on the basis that a bid is validation. Each is defensible on its own terms. What they share is a direction. All five enumerate ground where commercial attention has already been paid, which is the same ground on which you and your competitors already publish.
The sign of that bias has been measured, by the company whose keyword database supplies most of these baskets. In an analysis published on 7 April 2026, Semrush cross-referenced more than a billion lines of United States clickstream data covering October 2024 to February 2026 against its own database of over 27 billion keywords. For most of the period, between 65 and 85 per cent of ChatGPT prompts matched no traditional search keyword at all.
The second finding is the one that fixes the sign. Among the prompts that did match a traditional search term, most carried navigational or transactional intent. The matchable slice of the prompt population is disproportionately made up of people who already know what they want and who they want it from. A basket built from keyword data therefore over-samples buyers who already know your name, which is the definition of a positive defect correlation, and it does so before anyone runs a single query.
INSTRUMENT 2 — THE FRAME LEDGER
| Prompt source | What it actually enumerates | Direction of the defect | What it may be used for |
| Keyword export from an SEO tool | Phrases typed into a search box and captured by a crawler-fed database | Strongly upward. Over-samples navigational and branded language | A labelled stratum, never the whole basket |
| Competitors’ paid keywords | Terms somebody already chose to bid on | Upward. Enumerates a market’s settled vocabulary, not its open questions | Competitive strata only |
| People Also Ask and AI Overview mining | Questions an engine already answers from a page-based surface | Upward and circular. Your own presence helped generate the sample | Diagnostics, not levels |
| Community and forum mining | Questions from people who chose a public venue to ask them | Mixed. Skews toward problems, failures and complaints | A useful counterweight to the rows above |
| Tender clarifications, sales calls, support tickets, regulator and trade-body pages | Questions put in writing to a third party by named buyers | Roughly neutral for the buying stage it covers. Under-samples the earliest stage | The independent half of a designed basket |
The green row is not a better class of prompt, only a differently sourced one: independence is the only property that matters here. Community mining sits in amber rather than red for the same reason that earning links on Hacker News differs from posting on your own blog. The venue was not chosen by you.
One complication no sourcing discipline can solve: in the same Semrush sample, the most common prompt of all was a request to describe the user from their previous chats, appearing 21,576 times across 21,050 users. ChatGPT had suggested it to them. A barbecue invitation in the same top ten was likewise pre-loaded into the interface. Part of the prompt population is authored by the platform, and no buyer research anticipates it.
Key takeaway. Rho is not a nuisance parameter with an unknown sign. For any basket built out of keyword data, it is positive, it has been measured indirectly by the keyword vendors themselves, and it inflates your score before the first query is run.
The same defect that inflates the score also blinds the instrument
Overstatement is the consequence everyone would guess, and the harmless one: a level uniformly too high still ranks strata correctly.
The second consequence is not harmless. If your basket over-represents ground you already hold, then the prompts you would have to win in order to grow are disproportionately the prompts you never thought to include. The instrument’s response surface is truncated. Work that closes a real gap registers weakly or not at all, and work that polishes an existing strength registers cleanly, which is precisely the wrong incentive to hand a team with a quarterly target.
Overstates the level. Understates the change. One defect, two faces.
The arithmetic is simple enough to do in a meeting. If only a fraction of your basket contains prompts that the planned work could plausibly move, then the measured lift is roughly the true lift multiplied by that fraction. Call it the detectable share. At a detectable share of 0.2, a programme that genuinely improves the questions it targets by 30 points shows up as six.
THE SENSITIVITY RULE
Before you commission the work, go through the basket prompt by prompt and ask one question: if this programme succeeds exactly as planned, does the answer to this prompt change? Count the yeses and divide by the basket. That fraction is the share of your own improvement the instrument is capable of seeing. If it comes in under a third, redesign the basket before you spend a pound on the programme. A basket that flatters you is a basket that cannot show your progress.
It also forces a team to state in advance which questions the work should affect, whatever the tooling reports.
Diagnosing a frame defect without a census
INSTRUMENT 3 — THE SPLIT-SOURCE CHECK
- 1. Leave the main basket exactly as it is. Do not clean it, do not extend it.
- 2. Build a second basket of 40 to 80 prompts entirely from sources you neither control nor wrote: verbatim tender clarification questions, procurement portal exchanges, regulator and trade-body guidance pages, support tickets, and questions transcribed from recorded sales calls.
- 3. Run both baskets on the same engines, at the same run count, in the same week.
- 4. Compare the two headline readings. Because you equalised everything else, the gap is an estimate of frame error, not of noise.
- 5. Set N at your honest estimate of the category’s question count and read rho off the effective basket size formula, or simply record the gap and repeat quarterly.
Reading the gap. Inside your reported interval, the basket is behaving. Two to five times the interval, the level is directional only. An order of magnitude, and the level is not identified: report strata and movement, and stop printing the headline.
What is a split-source check?
A split-source check is a paired measurement in which the same brand, engines and run count are applied to two prompt baskets built from deliberately unrelated sources. Sampling error is held constant by construction, so the difference between the two readings isolates the effect of how the prompts were chosen. It is the cheapest available estimate of the error term that no amount of extra running can reduce.
Sixty prompts across four engines at five runs is 1,200 calls: negligible in API cost, a day of analyst time to build. That is the whole price of finding out whether the subscription you renew monthly is measuring anything, against the cost of discovering it later through a citation recovery exercise.
Allocating a fixed budget across prompts, engines and runs
Every audit has a call budget per wave, and it is the product of three numbers: prompts, engines and runs. The interesting question is not how large the product is but how it is factorised, because the three factors buy different things.
Runs buy precision and nothing else
Extra runs of a prompt you already have reduce the uncertainty on that prompt’s own rate. They tell you nothing new about the category, cannot reach a question you did not ask, and their marginal value falls away fast. Runs are the term the field has been optimising, and the term with the hardest ceiling.
Prompts buy precision and coverage, but only if they are new in kind
Two hundred extra prompts pulled from the same export you used the first time buy precision alone. They leave rho exactly where it was, and by narrowing the interval around a biased level they make the number look more trustworthy than it was yesterday. That is the big data paradox arriving in a marketing dashboard. Two hundred extra prompts from an independent source buy coverage, which is the only purchase that touches the second error term.
Engines are strata, not replicates
Averaging across engines with equal weight asserts that your buyers use them equally, which is almost never true and almost never stated. Foglift found that 61.7 per cent of the top 25 cited domains appeared on exactly one engine; Superlines put citation rates at 27.01 per cent on Grok, 13.05 on Perplexity, 0.59 on ChatGPT and zero on Claude. Those are not noisy measurements of one quantity but different quantities. Weight engines by your own referral mix and report per engine regardless, remembering that agentic browsing and AI browsers are not the chat surface.
Weights must come from outside the audit
Uniform weighting claims that every question is worth the same to the business, which nobody believes and few teams examine. The alternative is not a better guess but an external source: tendered value by category, deal frequency by segment, won-business mix. The test is whether the weight source existed before the audit did. Publish the weights beside the score, because a weighted score without them is not reproducible by anyone, including you in six months.
The allocation rule that follows is a hundred years old and comes out of stratified sampling: put the marginal call into the stratum with the largest weight and the thinnest coverage. Sample more where the stratum matters more and where you know least. Prompts are not exchangeable, which is what makes stratification necessary rather than decorative. A University of St Gallen team measuring visibility across four verticals in early 2026 found per-prompt similarity within a single campaign ranging from above 0.8 to below 0.2, which is a spread no single headline average can survive.
Worked, on 4,000 calls a wave. Design A: 200 prompts, four engines, five runs. Design B: 250 prompts, four engines, four runs. Design C: 2,400 calls on 200 prompts at three runs, and the remaining 1,600 on a 100-prompt independent extension at four runs. A and B differ only in the width of an interval. C is the only one that buys any information about whether the basket points at the right questions.
What a statistical agency does when its frame breaks
In November 2023 the Office for Statistics Regulation removed accreditation from the Office for National Statistics’ Labour Force Survey estimates after response rates fell. In October 2024 the Annual Population Survey accreditation was suspended too. The loss propagated: fourteen statistical outputs were affected across the ONS, the Welsh Government, the Scottish Government and the Department for Business and Trade. Everyone downstream who had built on the number inherited the problem, which is worth remembering the next time an audit figure goes into a board pack and a bonus scheme in the same week.
The ONS then did the thing you would expect. It worked the response rate, and by mid-2026 reported Labour Force Survey responses back at broadly pre-pandemic levels. And on 11 August 2026 it wrote to the regulator to say it would not be seeking reaccreditation for Labour Force Survey and Annual Population Survey outputs at all. The responses came back. The badge did not.
It is replacing the design instead, with the online-first Transformed Labour Force Survey: a first readiness assessment in July 2026 that concluded the transition was not ready, a progress check in January 2027, a further assessment in July 2027, and a possible move of headline statistics in November 2027. One detail is worth carrying across. Running the old and new surveys in parallel degrades the quality of both and carries costs the ONS says will need managing.
The lesson. A national statistical institute would rather retire an instrument than publish a number whose frame error it cannot bound, and it has budgeted years and a parallel run for the changeover. Changing a basket honestly has a price. Not changing it has a larger one, paid quietly.
Worked example: an inspection firm that was measuring the wrong questions
Garrowby Testing and Inspection is a UKAS-accredited materials testing and weld inspection business in Malton, North Yorkshire, with 96 staff and £13.4M of revenue. Its buying process is tender-led: pre-qualification questionnaires, framework panels, and written clarification rounds in which named buyers ask specific technical questions in writing.
In January 2026 its agency began an AI visibility audit: 300 prompts built from the firm’s keyword export plus competitors’ paid terms, four engines, five runs, fortnightly. Around 12,000 captures a month on a £1,850 subscription. The headline mention rate came in at 41.2 per cent, with a reported 95 per cent interval of plus or minus 1.4 points.
In March the head of bids ran a split-source check out of scepticism. The second basket was 60 prompts: 41 verbatim clarification questions submitted by buyers in 2025 tender rounds, 12 from two trade-body guidance pages, seven transcribed from 30 recorded sales calls. Same engines, same runs, same week, for £95 of API spend and a day of analyst time.
The second basket read 12.4 per cent.
A gap of 28.8 points, against reported intervals of 1.4 and 3.0 points. The difference attributable to how the prompts were chosen was roughly ten times the largest uncertainty either instrument had admitted to. Taking N at 25,000 plausible category questions and rho at 0.05, the effective basket size of the 300-prompt instrument was about five, and the honest interval on its level, once clipped to lie between zero and one, said nothing whatsoever.
Between April and July the basket was rebuilt: 220 prompts, 140 retained and 80 from the independent sources, stratified into six buying stages and weighted by 2025 tendered value. Runs dropped to four, so a wave cost 3,520 calls, less than before. It read 33.8 per cent unweighted and 19.1 per cent value-weighted. The old basket, still running alongside, read 41.6 per cent.
The 80 independent prompts then did their second job. Resolved one by one into which document would have to exist for the answer to name Garrowby, they produced a list of 23: trade-body guidance naming a test method, four specification comparisons on sector titles, three procurement portal entries, five technical profiles on supplier registers, and a handful of comparison placements. By August, 14 existed.
The value-weighted reading moved from 19.1 per cent to 31.4 per cent. Over the same window the original basket moved from 41.6 to 44.0. The instrument the firm had been paying for all year registered about a fifth of the movement the designed one did, because the work had happened in the region it had never sampled.
What the redesign did not fix. Two of the 23 documents never happened and one trade body declined to name suppliers at all, so a stratum stayed at zero for the year. The independent basket is itself a construction: tender clarifications come from buyers who have already drawn a shortlist, so it under-samples the earliest stage and Garrowby does not claim otherwise. The weights were 2025 tendered value and were stale by design, and one large framework win in June would have re-weighted everything. Running both baskets in parallel for two quarters cost around £3,700 in subscription they would otherwise have cancelled. And one engine changed its default retrieval behaviour in May, moving both readings at once; separating that from the programme is a question about movement over time, and a different question from this one.
The objection this argument has to survive
The strongest counter is not that the arithmetic is wrong. It is that the arithmetic is answering a question nobody asked.
Nobody, the objection runs, claims a prompt basket is a probability sample. It is a key performance indicator. Freeze the wording, run it weekly, read the change and ignore the level: if the basket overstates by 20 points in January and 20 in June, the difference is clean, because constant bias cancels exactly in a difference. On that reading, the advice to freeze the prompt set is the correct response to a frame you cannot fix, and this whole argument attacks a claim nobody makes.
Formally, that is right, and it should be conceded without hedging. If rho is stable, level bias differences out exactly, and a frozen convenience basket is a legitimate control chart for internal trend. It should also be conceded that design guarantees nothing: the Labour Force Survey was a properly designed probability sample and lost its badge anyway. Three things bound the defence.
First, the insensitivity does not difference out. The blind region is fixed by construction, so the work most likely to close a real gap is the work a frozen instrument is least able to see. Differencing removes a level. It cannot restore a region that was never sampled. Garrowby’s 2.4 points against 12.3 is not a level bias failing to cancel; it is a truncated response surface.
Second, rho is not stable, and there is evidence rather than theory on this point. In the same Semrush clickstream sample, the share of ChatGPT prompts written in traditional search language nearly doubled between October 2025 and February 2026, from 18.9 per cent to 34.9 per cent. The relationship between a keyword-derived basket and the live prompt population changed by half in five months, on the platform’s clock rather than yours. A bias that drifts, or that correlates with the very intervention you are measuring, is exactly the case in which differencing fails.
Third, a frozen basket answers no cross-sectional question. Competitor share, category coverage, whether you lead a segment: each needs a frame shared with whoever you are compared against, and a basket built from your own keyword history cannot supply one.
The resolution: freeze the wording, schedule the re-sourcing. Hold prompt text fixed inside a reporting period so that the change you read is a change in the world. Refresh a defined fraction of the basket each period, a fifth is a sensible starting point, from sources you do not control. Run the overlap so the level is bridged rather than jumped, which is what the ONS is spending two years and a parallel run doing. Report the level with its frame stated, and allocate budget against strata rather than against the headline.
What this means for earned coverage
A properly sourced basket produces a prospect specification for free. Every prompt drawn from a tender clarification, a regulator’s guidance page or a trade body is, by construction, a question whose answer lives in a document you do not own. Resolving each one into the question of which document would have to exist turns improve our AI visibility into a finite list with named publishers attached. Garrowby’s version of that list had 23 entries. One exercise, two outputs: a defensible instrument and a costed brief for the strategies that produce those documents.
Value-weighting makes the coverage programme affordable. Uniform weights spread the requirement evenly across every question, which is both unaffordable and unfalsifiable. Weighted properly, the requirement collapses onto a few high-value strata and a small number of documents, which is a brief a specialist can actually execute against inside a quarter.
And the asymmetry in the instrument mirrors the asymmetry in the market. In measurement, precision is the term you can buy and the frame is the term you cannot, and it is the unbuyable term that carries the information. The same split runs through the thing being measured. Work on your own pages is the term you can buy: more pages, deeper coverage, cleaner markup, all on your own schedule, whether that is your site or public pages you own elsewhere. Whether a third party states your claim in their own document, on their own masthead, is the term you cannot buy. That is precisely why it is worth something to an engine assembling an answer, and why it survives compression into a recommendation or a multi-turn session when your own pages do not.
A basket built only from what you control measures an estate you already own. That estate is the part of the problem that was never scarce, which is the deepest reason a keyword-derived audit reads high and moves little. It is also why what a link is for has not changed as much as the measurement conversation implies, even as the underlying statistics have moved considerably.
Two limitations on the evidence used here. The Semrush clickstream panel is United States sessions, so a British basket needs its own version of the exercise before those percentages are treated as local, which is a general caution about cross-market work. And the frame gap is measured against a keyword database that is itself a construction, growing whenever a keyword starts triggering an AI answer.
The Monday checklist
- Pull your current basket into a sheet and write beside each prompt where it came from. If every row says keyword export, you have one source and no independence.
- Run the sensitivity rule. Mark each prompt yes or no for whether your planned work would change its answer, and compute the detectable share.
- Build a 40 to 80 prompt independent basket this week from tender clarifications, support tickets, recorded sales calls, and regulator and trade-body pages.
- Run the split-source check: same engines, same run count, same week, both baskets.
- Set N at your honest estimate of the category’s question count, solve for rho, and write the effective basket size into the report next to the headline.
- Cut the run count by one and spend the saving on independent prompts.
- Replace uniform weights with a weight source that existed before the audit, and publish the weights beside the score.
- Report per engine, always, and never average across engines without stating the weights.
- Diarise the re-sourcing: a fifth of the basket every quarter, with an overlap wave so the level is bridged.
- Turn every independent prompt you lose into a named document and a named publisher, and put those on the outreach list. Check capture mechanics too: a page that renders late can be absent for reasons unrelated to your argument, as anyone who has debugged JavaScript and crawling will recognise.
The field spent 2026 arguing about how many times to ask. The more expensive question was always which questions, and it is the one a competitor cannot answer for you. Documenting how the basket was built, in the way a publisher documents provenance, turns a score into evidence. And if the basket is to reflect where models learn about your category rather than where you already rank, the sourcing question and the training-source question are the same question in different clothes.
