Benchmarking Citation Share

Benchmarking Citation Share vs Competitors: A Reproducible Method

TL;DR

•  Every control in the standard 2026 AI-visibility audit — incognito, signed out, cookies cleared, run it twice — is a repeatability control. It makes your runs agree with your own runs. None of them makes your number agree with anybody else’s.

•  That gap has a name and a number. ISO 5725 calls it the reproducibility limit: how far two honest measurers may differ before a difference means anything. You can compute yours with one extra run a quarter.

•  Most tracking tools cannot resolve the 1–3 point margins that decide category leadership. A 46-point account-personalisation effect has now been measured sitting on top of them.

•  The only series in the whole benchmark that two independent operators will count identically is the corroboration record — the third-party documents that name you. That is where the auditable number lives.

The word the field is using was redefined in 2020

The AI-visibility audit guides published through 2026 converge on almost identical advice, and it is good advice. Open each engine in a separate private window. Sign out wherever the engine allows it. Clear cookies between sessions. Build 30 to 50 prompts, screenshot every response, and run each prompt at least twice on different days, because a single answer is a sample and two answers are a pattern. One widely-circulated guide adds that this is the step most teams skip, and the one that decides whether the audit is useful or theatre.

It is competent advice, sincerely given. It also cannot deliver what its own headline promises, for a reason that predates AI search by six years.

The National Academies settled the terms in 2019. Reproducibility is obtaining consistent results using the same input data, the same computational steps and code, and the same conditions of analysis. Replicability is obtaining consistent results across studies aimed at the same question, each of which has obtained its own data. The distinction is not pedantry, and it is genuinely treacherous: before 2020 the Association for Computing Machinery defined the two words the other way round, then swapped them to match. Anyone who reads “a reproducible method” and pictures somebody else re-running the prompts and landing on the same score is picturing replication — the harder of the two, and the one a commercial engine will never supply.

What is a reproducible citation-share benchmark?

It is one where a second party, holding your prompt basket, your parsing rules and your retained raw responses, can recompute your published number exactly. It is not one where a second party re-running your prompts sees what you saw. The first is achievable this quarter, at almost no cost. The second is not achievable at all — and nearly every protocol on the market is built to chase it.

Every control in the canon is a repeatability control

ISO 5725 — Accuracy (trueness and precision) of measurement methods and results, whose Part 2 was reissued in 2025 — splits precision into two conditions. The split is the whole argument.

  • Repeatability conditions: same method, identical items, same laboratory, same operator, same equipment, short interval.
  • Reproducibility conditions: same method, identical items, different laboratories, operators and equipment.

Each condition has a published limit — r and R, both 2.8 times the relevant standard deviation — and each is read the same way: the amount within which two determinations should agree 95% of the time. Two results further apart than the limit are telling you something; two results closer are telling you nothing.

Now re-read the audit canon. Private window, signed out, cookies cleared, extensions off, two runs on different days — with one analyst, one laptop, one office network and one set of accounts. Every item on that list is a repeatability control. They are good controls and they tighten r. They do nothing whatsoever to R, because R is defined by variation between measurers and the protocol contains exactly one.

The standard carries a note that belongs on the wall of every team doing this work: an estimate of the repeatability standard deviation can be obtained from routine work inside one laboratory using control charts, but an estimate of the reproducibility standard deviation can be obtained only from a planned, organised study involving more than one. There is no volume of self-repetition that produces it. Running your own audit weekly instead of monthly buys precision on the small term and leaves the large one unmeasured.

Repeatable but not reproducible is a real state, not a hypothetical

Manufacturing has measured this failure for decades under the name Gage R&R, a study that partitions observed variation into part-to-part variation, repeatability (the equipment) and reproducibility (the appraiser). Instruments that are excellent on one and disastrous on the other turn up constantly. One published measurement comparison in dental research reported a traditional technique with repeatability of 0.58% and reproducibility of 33.01% — an instrument that agreed with itself almost perfectly and with a second operator not at all.

An instrument in that state produces smooth, internally consistent monthly reporting — the exact output a marketing team is trained to trust, and the exact output that collapses the first time an agency or an incoming CMO measures the same thing. Anyone who has run a backlink profile against a competitor’s knows the feeling of two tools disagreeing; the difference is that link tools disagree visibly, and citation-share tools disagree behind a single confident number.

KEY TAKEAWAY

A benchmark’s audience is always a second party — a board, a client, an agency under review, a rival disputing your claim. A number only one seat in the world can produce is an assertion with error bars drawn on it.

The apparatus is not the same machine twice

In metrology the instrument is a thing you own, keep in a cupboard and send away to be calibrated. Here the instrument is a commercial service, versioned by somebody else, and a growing share of what determines your reading is not merely uncontrolled but unobservable.

Take the least intuitive source first. In September 2025 Thinking Machines Lab published the clearest available account of why large models are non-deterministic even when you switch the randomness off. Sampling 1,000 completions from a large open model at temperature zero produced 80 distinct outputs, first diverging at token 103. The cause is not sampling — temperature zero removes that. It is that reduction kernels on a GPU behave differently depending on the size of the batch they execute in, and the batch your request lands in depends on how many other people happened to be querying that server in the same instant. From outside, the model is not a function of your prompt. It is a function of your prompt and the load. Batch-invariant kernels fix it at a throughput cost of roughly a third to two-thirds, depending on implementation — a price a research lab will pay and a consumer product never will.

Then the ordinary sources, which are larger. On 6 July 2026 OpenAI began routing users who exhaust their rate limits to a smaller fallback model that does not appear in the model picker; on 15 July 2026 it raised the custom-instructions ceiling from 1,500 to 5,000 characters of standing personalisation shaping every answer. Whether live retrieval fires at all is its own variable: the St Gallen team measuring AI visibility across four engines found ChatGPT suppressed web search on 57.8% of runs, which means a pooled score silently blends a retrieval measurement with a memory test of the model’s weights.

Apparatus variableCan you set it?What it does to your numberMinimum record per run
Prompt wording and orderYoursSets what is asked; small edits move rankingsExact strings, version-dated
Engine and product surfaceYoursDifferent index, different citation slateProduct and interface, not just vendor
Account state, memory, custom instructionsYoursThe largest single measured leverSigned in or out; instruction text or blank
Locale, IP and deviceYoursRegional slates and local entitiesCountry, network type, device class
Parsing and attribution ruleYoursDecides what counts as a mention at allRule version and date of last change
Whether live retrieval firedObservable onlySplits your run across two different machinesFlag on every response
Model version actually servedObservable onlySilent fallback routing on rate limitsReported model string per response
Index freshness and vendor configNeitherMoves the level with no notice to youVendor changelog date at time of run
Concurrent batch compositionNeitherNon-determinism even at temperature zeroTimestamp to the second

What can you actually control in an AI visibility audit?

The first five rows of that table and nothing below them. Everything under the line is recordable but not settable — which is why the honest artefact of a benchmark is a record, not a recipe. A recipe implies that following it reproduces the dish. A record only claims to say what was in the kitchen that day, which is a smaller claim and a true one.

Signing out is not a control, it is a choice of operating point

The canon’s answer to personalisation is to strip it out: sign out, clear the session, measure the model’s default behaviour rather than your own history. The logic is sound as far as it goes. But depersonalisation is not neutrality. It selects one operating point out of many, and it happens to be the one that fewest of your buyers occupy.

The size of the effect is now measured rather than assumed. In an experiment published on 21 May 2026, iPullRank ran 1,922 Google AI Mode responses across matched accounts between 30 March and 15 April 2026, producing 22,064 brand-level observations, and compared a blank control account against a blank account connected to Google’s personal-context features and seeded with brand signals in Gmail and Photos. On a difference-in-differences comparison, seeded brands were 46 percentage points more likely to appear than in the control. Email signals outperformed image signals, and newsletters that were never opened still surfaced as sources.

Hold that number against the decision it is being used to inform. The most carefully measured category study of 2026 — Semrush and Kevin Indig’s monthly tracking of 1,094 US categories inside ChatGPT — found that where category leadership changed hands the median lead was 1.3 points, and where leadership held it was 2.9. A 46-point account effect is sitting on top of a judgement being made on a 1.3-point margin. Whatever else is true, the account is a bigger lever than anything currently on your website.

The configuration belongs in the metric’s name

The fix is not to abandon signed-out measurement, which is the only condition you can standardise cheaply. The fix is to stop calling it the number. Name it in full: citation share, signed out, UK IP, desktop, retrieval-enabled runs only. Then the reader knows which machine produced it, and a second party can decide whether that machine resembles their buyer.

For a firm selling outage software to utilities, or compliance tooling into the NHS, that gap is not academic. Every buyer sits inside a managed workspace tenant with three years of organisational context behind their assistant. The signed-out benchmark measures a stranger with no history; the buyer is served by something closer to a colleague. A deep-research run is a different apparatus again. Both readings are legitimate and they are not interchangeable, which is precisely why the set of sources an engine assembles has to be reported against a stated configuration rather than in the abstract.

The number the field is missing: your reproducibility limit

Everything above is diagnosis. This is the part you can run next week, and it costs one extra pass of a basket you already own.

THE TWO-OPERATOR CHECK

Design. Freeze the basket and the parsing rule. On the same day, have two people run it once each. Operator A is your analyst: their machine, your office network, their usual account state. Operator B is genuinely separate — a different person, a different device, a different network, a different account. A contractor, a colleague in another office, or your agency partner all qualify. What must not be shared is the seat.

Arithmetic. For every brand you track, take the absolute difference between A and B in citation share. Average those differences. Your reproducibility limit is approximately 2.5 times that average. (The exact route: divide the standard deviation of the paired differences by the square root of two to get the reproducibility standard deviation, then multiply by 2.8. Under normal conditions the two routes agree.)

Reading. R is the distance two honest measurers may differ by pure chance 95% of the time. Any competitor gap, and any quarter-on-quarter movement, smaller than R is not evidence of anything.

•  Movement below R: report as unchanged, with the limit printed alongside.

•  Movement between R and 2R: report as a change, with the limit stated in the same sentence.

•  Movement above 2R: treat as a finding — and interrogate the apparatus record before the marketing.

How many competitive positions can your tool actually see?

There is a second output from the same data, and it is the more damning of the two. Measurement systems analysis asks how many distinct groups an instrument can separate: the number of distinct categories, calculated as 1.41 times the ratio of part variation to measurement variation, truncated to a whole number. The published thresholds are unambiguous. Five or more, and the instrument supports tracking and control. Two to four, and it supports classification only — you can sort leaders from laggards and nothing finer. One, and it cannot tell two parts apart. The equivalent statement in percentage terms puts a measurement system below 10% of total variation as acceptable, 10–30% as conditional, and 30% or above as unusable.

Translate the terms. Part variation is the spread of citation share across the brands you track. Measurement variation is the operator-to-operator spread you just computed. If your tracked brands span nine points and your two operators differ by three, the count is four: the instrument can tell a leader from a laggard, and cannot see a quarter’s work.

What this looks like on a real basket

Ravensdale Grid Software, in Nottingham, has 41 staff and £6.2m of ARR selling outage-management software to UK distribution network operators. It pays £1,600 a month for a visibility platform running a 140-prompt basket across five engines. April’s report put Ravensdale on 21.4% share of mentions against Culmore Systems on 25.9%, down 2.1 points on the quarter and flagged amber in the board pack.

On 22 April the head of marketing ran the two-operator check. The in-house analyst worked from the office network on her usual signed-in account; a contractor in Manchester used a fresh account on a domestic connection. Same 140 prompts, same day, same parsing rule. Across the six tracked brands the absolute differences came out at 6.6, 4.4, 3.2, 2.7, 1.9 and 0.8 points — an average of 3.3, giving a reproducibility limit of roughly 8 points. The brands themselves spanned about 9.4 points, so the distinct-category count was four.

Two things followed within a week. The 2.1-point quarterly decline was a quarter of the limit and had to be restated as unchanged, which cost an uncomfortable meeting and saved a reallocation. And the gap to Culmore was 4.5 points in one operator’s run and 13.8 in the other’s — both could not be right, and the disagreement tracked account state rather than the market. The instrument had been reporting to two decimal places a quantity it could not resolve to eight.

What you can make reproducible is the analysis, not the reading

The National Academies anticipated this exact situation. Their recommendation on computational reproducibility asks researchers to convey the input data, the study methods and the computational environment — and then, specifically, the intermediate results and output data “for steps that are nondeterministic and cannot be reproduced in principle”. That clause is the entire operating instruction for AI-visibility work. Where a step cannot be re-run, you retain its output. The obligation does not disappear because the machine is stochastic; it moves.

Machine-learning evaluation arrived at the same place from the opposite direction, and its experience should end the argument. This is a field with total control of its apparatus: open weights, its own hardware, fixed seeds, published harnesses, no vendor in the loop. It still cannot make two teams agree. A study across 20 models, 39 tasks and 6.5 million prompt variants found performance rankings frequently reversed under semantically equivalent prompts. Formatting alone has moved accuracy by tens of points, changing a random seed shifts scores by 5 to 15 points on mathematics benchmarks, and running identical weights through two evaluation harnesses can swing a headline number by 10 to 20 points. The field’s own remedy is a reporting checklist whose central ask is to publish the exact prompts, the evaluation code, and the model outputs.

THE RECOMPUTE TEST

The question: can you regenerate last quarter’s headline number, exactly, from files you already hold — with no calls to any engine?

If yes, your analysis is reproducible in the strict sense. Any dispute is now a dispute about method, which is arguable in a room, in front of documents, by people who disagree.

If no, you do not have a benchmark. You have a memory of one. It cannot be checked, corrected or defended, and its first serious challenge will end it.

The rule that follows: the evidence must outlive the claim. A score whose underlying responses have been discarded is not wrong — it is unfalsifiable, which is worse, because nothing will ever correct it.

Most platforms in this market sell a score and keep the response text briefly or not at all, because storing full answer text across thousands of prompt-engine-day combinations is the expensive part of the product. Three asks belong in the next renewal, in writing: raw response export including full answer text and citation URLs; per-run metadata covering the model string served, the retrieval flag, the timestamp, the locale and the account condition; and a retention window at least as long as the period you intend to report on. If the only export is a rendered dashboard rather than structured, machine-readable records, the data is theirs and not yours. Note which of the three they answer. The pattern of what a vendor will commit to in writing is itself a signal about how the numbers were produced.

KEY TAKEAWAY

You cannot make an engine give the same answer twice. You can make your own number checkable by a stranger. Only one of those is inside your budget — and only one of them is what the word reproducible actually means.

The protocol

THE BENCHMARK PROTOCOL

1.  Freeze it in writing. Prompt strings, engine list, brand list, parsing and attribution rules — dated, version-numbered, and changed only at a scheduled review rather than mid-quarter.

2.  Declare the operating point. Signed in or out, locale, device class, retrieval setting. It goes in the metric’s name, not a footnote.

3.  Run the two-operator check once a quarter. Second person, second device, second network, second account, same day, same basket. Compute R.

4.  Log the apparatus record on every run and stratify by whether retrieval fired. Report retrieval-on and retrieval-off separately; pooling them averages two different machines.

5.  Retain raw responses, not scores. Full answer text, citation URLs, timestamps and model strings. Hash the file set so a later dispute is about interpretation rather than provenance.

6.  Report contrasts against R. Gaps and movements inside the limit are reported as unchanged, with the limit printed next to them in the same table.

7.  Re-run R after any change of apparatus — a vendor methodology update, a model deprecation, a new engine added to the basket, or a change in your own account conditions.

8.  Hand it to someone who did not build it and ask them to restate your headline number from the files alone. Whatever they cannot do is your actual gap.

What the check cannot do

Three limits, stated plainly. Two operators is a small study, and an R computed from six pairs is itself imprecise — report it as an order of magnitude, not to a decimal. R lumps operator, device, network and account into a single term, so it sizes the problem without diagnosing it; splitting it means varying one factor at a time, which costs runs. And there is a structural ceiling no in-house team escapes.

In proficiency testing the assigned value is a consensus across independent participants, and the tolerance is set in advance by an organiser who is not one of them — the framework where a z-score above two is a warning signal and above three an action signal. A consensus of two, both paid by the same company, is a useful sanity check and not an accreditation. That is also the correct reading of the most-shared measurement embarrassment of 2026, in which two platforms reported 8% and 40% visibility for the same twelve-person accounting firm in the same month. That is not a scandal. It is an unaccredited inter-laboratory comparison with no assigned value and no tolerance, in which both participants remain unfalsifiable.

Metrology’s answer where no reference standard exists is not to give up and not to pretend. Where direct traceability is impossible, the accepted substitute is an organised comparison with a published consensus value. The AI-visibility market has produced more than twenty competing tools and no comparison. Until somebody organises one, two operators inside your own company is the best available approximation.

Where the argument survives its own strongest objection

The strongest counter is not that any of this is wrong. It is that it is beside the point. It runs like this: nobody needs metrology to know whether a rival gets mentioned more often than they do. A fixed basket run monthly gives you direction, direction is all anyone acts on, and demanding a two-operator study every quarter is exactly how measurement budgets get cut. Every practical trade in the world runs on uncalibrated instruments that are good enough for the decision at hand. A builder’s tape measure has no traceability chain and the houses stand up, and this field ran happily for a decade on metrics nobody could audit.

Most of that is right, and the concession is real: cheap and directional beats expensive and absent, and the proposal here is one extra run a quarter precisely because an accreditation scheme would never be adopted. But the defence fails on three counts.

  • Good enough for the decision is a claim about a specific decision. The decisions being taken sit inside the noise. Category leadership turns on margins of 1.3 to 2.9 points; a typical reproducibility limit is larger than both. The directional defence fails at exactly the case it is invoked for.
  • Direction is not preserved across configurations. A tape that reads two millimetres long reads two millimetres long for everyone, so the error cancels in a comparison. An engine’s answer can invert between a signed-out and a signed-in account. A common offset cancels; a configuration-dependent one does not.
  • Cheap and directional stops being cheap the moment the number allocates. Once it moves budget between two agencies or decides a retainer, it is no longer a communication device. It is an allocation device, and the first person to ask how you know gets an answer no second party can check.

The one series two strangers will count the same way

Everything above is a story about an instrument you do not own. One quantity in the whole exercise behaves differently, and it is the oldest one on the list.

The corroboration record — the set of third-party documents that name you in your category — has a reproducibility limit of approximately zero. Two operators, on two accounts, on two continents, counting the same set of documents arrive at the same count, because the object being counted is a document with an address rather than a response conditional on a session. It can be listed, dated, archived and handed to somebody who actively distrusts you. That is not a property of what a backlink is that anyone designed for this purpose; it is a side effect of the web being made of addressable documents, and in 2026 it is the most valuable measurement property you have access to.

Three consequences follow. For the benchmark: pair every measured citation with the document behind it and keep both series. The score is the volatile one; the document set is the auditable one. When the score moves and the document set does not, you are looking at the apparatus. When the document set moves and the score follows some weeks later, you have something worth putting in front of a board — which is the same asymmetry that makes citations inside AI Overviews easier to argue about than the traffic they do or do not send, and the reason counting acquisition over time survives instrumentation changes that destroy a score series.

For the reporting: a benchmark built on documents can be shown to the competitor, and one built on scores cannot survive the first request for the working. If you have watched a visibility claim collapse under scrutiny, it collapsed at that exact question.

For how you buy placements, a new prospecting screen falls out of the same property, and it is the most useful thing in this article:

THE SECOND-PARTY SCREEN

Ask of every placement: can somebody who does not trust me, and does not have my credentials, verify that this exists?

Downgraded: member-gated trade press, region-locked pages, content that only exists after JavaScript executes, syndicated copies that rotate through partner domains, anything living behind an account or a login.

Upgraded: stable canonical URLs on institutional domains, public registers, regulator and trade-body pages, and anything a third-party archive has already captured with a date.

Why it reorders a list: verifiability is a property of publishing infrastructure, so it is close to uncorrelated with the authority scores every prospect list is already sorted by.

That screen has teeth on the technical side too. A placement rendered only after client-side JavaScript executes may be invisible to a verifier working from source, and a public page on a hosted workspace tool can be re-parented or unpublished by its owner without notice. By contrast, a thread on a public technical forum is archived by parties who have no relationship with you, which is exactly the property you want. The same logic runs through the infrastructure side of link acquisition: what is durable is what a stranger can fetch, and a dated, non-repudiable record of a placement is a different asset from a screenshot.

Ravensdale’s second half makes the contrast concrete. Over the two quarters to July the platform’s score moved 21.4, then 20.6, then 22.1 — a range comfortably inside its own eight-point limit and reportable in neither direction. The document set went 34, then 41, then 47, and both operators counted 47. Of those, 39 passed the second-party screen; the eight failures sat behind one gated utilities trade title, which the team kept for its readership and stopped counting as evidence. The July renewal conversation ran on the 47, not on the 22.1.

One caution, because this is the point at which a good argument overreaches. The document set measures supply, not selection. It tells you what evidence about you exists in the world, not what any engine chose to do with it, and the role a source plays in an assembled answer is not decided by count. It is a leading indicator and an auditable one. It is not a replacement for measuring the answers, and any consultant who tells you otherwise is selling the easier half.

The Monday checklist

  1. Print your basket, parsing rule and brand list into one dated document this week. If it lives in a tool’s interface rather than a file, you cannot prove what it said last quarter.
  2. Book the two-operator check for the last working day of this month. Name the second operator now — a contractor, another office, your agency — and confirm they are on a different account and network.
  3. Compute R and put it in the report header, next to the score, permanently.
  4. Re-read the last four board slides and mark every movement smaller than R. Decide what, if anything, was actually decided on those numbers.
  5. Add the apparatus fields to your logging: model string, retrieval flag, timestamp, locale, account condition. Stratify the next report by whether retrieval fired.
  6. Send your vendor the three asks in writing — raw response export, per-run metadata, retention window — and record which ones they answer.
  7. Run the recompute test on last quarter’s headline number. Give yourself an afternoon; most teams discover the answer is no.
  8. Start the document ledger: every third-party page naming you in the category, with date, URL and a yes or no on the second-party screen. It takes an afternoon and it is the only series in this article with a reproducibility limit of zero.
  9. Re-screen next quarter’s prospect list on the credentials question before you screen it on anything else.

The field spent 2026 trying to make an engine answer the same way twice, which is not available at any price, and neglected the part that was free. An engine’s answer is a reading. A placement is a record. Only one of them will still be there when somebody asks you to prove it.

Leave a Reply

Your email address will not be published. Required fields are marked *

Share of Model Previous post Share of Model: Measuring Category Dominance Across Five Engines
AI Citation Attribution Pipeline Next post From Impressions to Influence: Attributing Pipeline to AI Citations