AI Visibility Metrics

Standardising AI Visibility Metrics: Toward a 2027 Industry Benchmark

TL;DR

  • On 3 August 2026 the IAB published Measuring Visibility in the AI Era, the first industry framework for AI visibility measurement. It defines no number — only what a provider must disclose and when their data is fit to spend money against.
  • That is correct, for a reason the framework leaves unstated: AI visibility has no sampling frame. Every score on sale is an index defined by its own prompt list, not an estimate of a population quantity.
  • You cannot fix that with rigour. A frame is not a methodology, and the population of prompts is unbounded, private and partly generated by the engine itself.
  • Currencies are convened by whoever bears the loss when the number is wrong, and the framework scopes paid placements out — so the ad in the answer gets audited long before the citation beside it.
  • Three instruments below: THE FOUR FRAMES, THE COMPARABILITY TEST and THE CONTRACT CLAUSE.
  • The payload for anyone earning links: the third-party documents carrying your claim are the only quantity here with a census frame — countable, dated, verifiable by a stranger.

The standard everyone asked for arrived on 3 August 2026, and it is not a metric

For two years the complaint has been identical in every client meeting: there is no standard for measuring AI visibility, so nobody can say whether the work is working. On 3 August 2026 the Interactive Advertising Bureau answered it. Measuring Visibility in the AI Era was published as part of Project Eidos, the IAB’s programme for how industry measurement is developed and compared, and was built by a working group led by Caroline Giegerich with members from Acxiom, Walmart, WPP Media, Microsoft, PMG, Tinuiti, EMARKETER and the Alliance for Audited Media.

The framework starts from the same complaint. More than 20 companies now sell AI visibility measurement, each with its own methodology, and two can return materially different answers for the same brand. Only 16% of brands track AI visibility systematically. Buyers, the IAB notes drily, have budgets ready and no basis for evaluating what they are buying.

Then it does something almost nobody expected. It declines to issue a number, and it is right to.

What is the IAB’s AI visibility framework?

It is a shared vocabulary, a two-tier data-quality standard and a list of disclosures providers must make — not a metric, a score or a benchmark. It tells you how to judge whether somebody else’s number is fit for the decision you are making with it.

What is actually in it

Four categories, arranged as a causal hierarchy the IAB calls the 4 P’s. Presence (mention rate, citation rate, share of voice, visibility momentum) asks whether you appear. Prominence (position) asks where. Portrayal (sentiment, framing, hallucination rate, factual inaccuracy rate) asks in what context and with what accuracy. Persuasion (recommendation strength, post-citation click-through) asks whether any of it moves anyone. Publishers get a parallel set built on citation rate, content utilisation rate, attribution clarity and citation decay rate.

Above that sits a quality split: directional measurement, good for spotting trends, and decision-grade measurement, the only tier fit for budget reallocation, agency reviews or executive strategy. A criteria matrix separates them on sample size, query volume, prompt-type coverage, cadence, reproducibility, validation and platform coverage. Fewer than 50 queries is classed as exploratory rather than even directional. Decision-grade requires all four intent types — informational, comparison, recommendation and transactional — weekly collection, per-platform reporting, and a defined range of acceptable variation inside a seven-day window.

Then a disclosure list: which platforms and model versions; how the prompt library was built, weighted and refreshed; whether queries came from real search data or were invented by the provider; whether live retrieval was on, off or both; and whether trend data has been re-baselined after a platform change. Where a provider will not answer, the IAB says “that absence should itself be treated as a signal”.

What is not in it

No threshold: the framework states plainly that quantitative thresholds for acceptable variability remain an open working-group question. No rating of providers, no prescribed tool, no certification — only a note that the disclosure list could found one at some future point. And no paid measurement at all: the scope is organic visibility, with paid placement flagged as an adjacent gap.

Most of the industry read the announcement as the arrival of a benchmark. It is not one. It is a rule about evidence, and it is the right rule — for a reason the document itself only gestures at.

A metric needs a measurand, a frame and an adjudicator — in that order

Metrology, the science of measurement, insists on a sequence marketing routinely skips. First you define the measurand, the quantity you intend to measure and the conditions under which it counts. Then you establish a sampling frame, the list of everything that could have been sampled and the odds each item had of being picked. Only then does method — how many runs, how many prompts — become a meaningful question. Underneath all three sits an adjudicator, whoever settles a dispute when two parties read the same number differently.

The AI visibility discourse of 2026 is almost entirely a method discourse: how many prompts, how many runs, blend the engines or split them. These are third-in-the-queue questions being asked as though the first two had been settled. They have not been.

The measurand problem: four quantities wearing one name

“AI visibility” is not a quantity. It is a label sitting on at least four, and they move independently. Whether a model names you in prose is a different event from whether it links your domain as a source, which differs again from your share of the category’s mentions, which differs again from whether anybody acted. The IAB’s own definitions concede it: a brand can carry a high mention rate and a low citation rate, and the gap is itself the finding.

This is why the 4 P’s read better as four separate measurands than as a hierarchy of value. Presence and prominence are properties of a generated text; portrayal is a classification judgement laid on top of it; persuasion is an event in your own analytics. Only one of the four is countable in your own systems, and it is the one no competitor can benchmark you against. Our work on how entity authority is actually measured runs into the same wall: the quantity is defined by the instrument that reports it.

A Manchester equipment hire firm watched its mention rate climb for three months while its citation rate flatlined. Read as one number, that is progress. Read as two, it is a model that had learned the brand’s name from third-party pages and no reason to open the brand’s own site — a different problem with a different fix, mapped in recovering lost AI citations.

THE FOUR FRAMES: the instrument that decides what your number can claim

Every marketing measurement rests on one of four kinds of frame, and the frame — not the sophistication of the arithmetic on top of it — decides what claims the number can support. Every metric argument downstream reduces to which row you are standing in.

Frame classWhat it rests onMarketing exampleWhat the number may claim
1. Delivery censusA log of every event, with the denominator held by the sellerAd impressions; ad-server and platform delivery logsAbsolute levels, auditable, comparable between buyers
2. Probability sampleAn enumerable population with known selection probabilitiesBARB, RAJAR and PAMCo panels; properly drawn surveysPopulation estimates with stated confidence intervals
3. Public-record enumerationA published record any third party can re-countBacklinks, cited documents, company filings, registersCounts and dates, verifiable by a stranger
4. Constructed listA list somebody invented, with no known population behind itEvery AI visibility prompt library on the marketChange within the same instrument only — never a level

Three of the four frames can support a currency. The fourth cannot, and every AI visibility score on sale lives in it. A prompt library is not a sample of anything: it is a list a vendor wrote, or that you wrote with a vendor, and no item on it had a known probability of selection, because no enumerated population exists to draw from.

Why this is not a maturity problem that time will fix

The reflex response is that the discipline is young and the frames will firm up. They will not, because the obstacle is structural rather than methodological. A probability sample of prompts would need the population of prompts. That population is unbounded, private to the platforms, personalised per user, and — the part the field keeps forgetting — partly written by the machine: query fan-out generates sub-queries no human ever typed, and those synthesised questions decide which sources get retrieved. You cannot sample a population the measured system is still writing.

The unit moved too. As we argued in our work on multi-turn query chains, buyers arrive through sessions rather than single prompts, so a prompt library measures the unit easiest to enumerate rather than the one that produces revenue.

KEY TAKEAWAY

A frame is not a methodology. No amount of rigour moves a number from frame 4 to frame 2, so no methodological standard — the IAB’s included — can make two providers’ scores comparable in absolute terms. Rigour makes a frame-4 number more stable; it cannot make it an estimate of anything.

Indexes are legitimate. Unpublished baskets are not

None of this makes a visibility score worthless. The Consumer Prices Index is a constructed basket too, and it runs monetary policy. An index works when the basket is fixed, published and change-controlled, so movement in the index means movement in the world rather than movement in the basket. The failure here is not that scores are indexes. It is that they are indexes read as estimates, with baskets that are private and revised without notice.

Two tools, one brand, 8% and 40% — and neither of them is broken

In July 2026 an agency opened trials of two AI visibility tools on the same day for the same client, a twelve-person accounting firm. One reported the brand in 8% of answers, the other 40%. That case, documented by Inspeccia, gets treated as a scandal about tool quality. It is two indexes with two baskets, behaving exactly as indexes do.

The substrate moves as well. SE Ranking ran 10,000 keywords through Google’s AI Mode three times in one day: across 122,617 links, average exact-URL overlap between passes was 9.2%, on 2,006 keywords not one URL matched, and six keywords out of ten thousand matched completely. AirOps found only 30% of brands stayed visible from one answer to the next, 20% across five runs. Anything built on that inherits the movement — which is why the IAB rules that single-response measurement is not measurement.

The disagreement is structural, not sloppy

Engine choice alone splits the answer. Superlines measured citation rates from 27.01% on Grok and 13.05% on Perplexity down to 0.59% on ChatGPT and zero on Claude. Foglift found 61.7% of top-25 cited domains appear in exactly one engine’s top-25, with ChatGPT and Perplexity correlating at 0.78 while each correlates with Google’s AI Overviews at only 0.54. MaxAEO records the same brand in half of ChatGPT’s answers and 15% of Perplexity’s for an identical question. Evidential quality does not vary fortyfold depending on who is reading; retrieval architecture does, which is why our analysis of how AI browsers change the surface and of what actually drives AI product recommendations keeps landing on per-engine answers.

Now put the 2026 benchmarks side by side. Pondral scored 200 brands across 8,215 results at a mean of 55.8 out of 100. Foglift put the broad-market median at 46 out of 100 across 311 domains. MaxAEO places the median non-branded mention rate near 31%. Three numbers for “the average brand,” irreconcilable, because they are three measurands on three baskets against three engine sets. Pooling them is category error with a chart attached.

THE COMPARABILITY TEST

Before you put two AI visibility numbers beside each other — two vendors, two quarters, two competitors, a benchmark and yourself — answer three questions.

  • Same basket? Was the prompt list identical, item for item, including weighting and intent mix?
  • Same rule? Does “mention” mean the same event — exact string, resolved entity, subsidiaries, misspellings, unlinked names?
  • Same conditions? Same engines and model versions, same retrieval setting, same geography and account state, same dates?

Any “no” means the numbers are not comparable and the difference between them is not a finding. Report them separately or not at all. The test disqualifies almost every published AI visibility benchmark, including the three above, which appear here as evidence of disagreement rather than as an average. Geography is the most common failure: one blended score across markets buries exactly the divergence that makes international link building a separate programme.

Why do two AI visibility tools give different scores?

Because each tool invents its own prompt list, its own definition of a mention, and its own engine mix, then reports the result as a percentage. The scores measure different things under different conditions, so the gap describes the instruments rather than the brand.

Who actually writes measurement standards, and what they standardise

Marketing has done this before, twice, and both precedents say the same thing.

In 1914 the Association of National Advertisers, with agencies and publishers, founded the Audit Bureau of Circulations because self-reported circulation figures could not be verified and advertisers were paying against them; it survives as the Alliance for Audited Media. In 1963 and 1964 the Harris Committee hearings in the US Congress examined the accuracy of broadcast audience research, concluded that self-regulation with independent audits beat legislation, and produced the Broadcast Rating Council — now the Media Rating Council — in 1964, with minimum standards effective from 31 March that year.

Two features of that model matter more than its existence: audits are performed by an independent accountancy firm and paid for by the company being measured, and any change of methodology triggers a fresh audit. That is what an adjudicator costs — a standing bill, borne by the measured party, for the privilege of a number other people will accept.

The distance between a guideline and an audit

The IAB and MRC finalised Attention Measurement Guidelines in November 2025 on more than 200 contributors’ input. As of late 2025 exactly one attention methodology held MRC accreditation, with a second vendor in audit. Guidelines precede accreditation by years, and accreditation arrives only where money already sits.

The AI visibility framework is at the guideline stage and knows it: certification is described as a possibility, not a plan. Two representatives of the Alliance for Audited Media sat on the working group, so the audit lineage is in the room — and has not yet been asked to write an invoice.

What a real currency costs, in British money

Origin, the UK’s advertiser-led cross-media measurement programme, is the closest live example. ISBA secured £11m for the build phase alone from more than 40 organisations, including 25-plus brand advertisers and six agency groups responsible for over 80% of UK annual media billings. By 2025 it was funded by 69 stakeholders — 45 brand owners, the six major holding groups, and Google, Meta, Amazon and TikTok — with advertisers committing to an ongoing Fractional Advertiser Contribution. On 31 July 2026 ISBA announced exclusive talks with Fifty5Blue, formerly Kantar Media, to invest as Origin becomes a standalone company.

Six years, eight figures, seventy organisations and outside investment — to standardise deduplicated reach and frequency, a quantity that already has a frame. Against that, AI visibility has no levy, no convening body, and a population that cannot be enumerated at any price. A joint-industry currency for AI citations is not arriving in 2027. That is arithmetic about who funds what, not pessimism.

KEY TAKEAWAY

Measurement standards are convened by whoever bears the loss when the number is wrong, and audits follow invoices. Nobody currently invoices for an earned AI citation, so the earned side of the answer will be governed by disclosure rather than by audit — which means the market will standardise on whichever vendor’s basket the most contracts happen to name.

The industry has already run this experiment

Link building never agreed an authority metric. Moz shipped Domain Authority, Ahrefs shipped Domain Rating, others followed, none were comparable, and the market wrote them into procurement anyway. When Moz rebuilt the metric as DA 2.0 on 5 March 2019, thousands of sites moved overnight, and Moz’s own principal search scientist advised checking competitors’ scores because they would move in the same direction. That advice is the whole doctrine in one line: the number is only interpretable against a peer set measured on the same instrument at the same moment. Page Authority was rebuilt the same way in September 2020.

Fifteen years later, agencies still lose fees over thresholds set against a vendor index its own maker called relative. AI visibility is on the identical path, only faster, and the people who lived through the first round are best placed to refuse the second — which is what our guidance on the tools themselves has always turned on: a tool confers perception, not truth.

The scope note that decides the next three years

The scope section holds the sentence that matters most to anyone whose work is earned rather than bought. The framework covers organic, non-paid visibility only, and identifies paid placement measurement as a gap — an urgent one, the IAB says, because a brand can now be cited organically while holding a paid placement in the same answer.

Read that against the four frames. A paid placement is a delivery event: served, logged, billed, denominator held by the seller. Frame 1. It has a counterparty who can withhold payment, an accreditation pipeline at the MRC, and a trade body auditing ad delivery since 1914. An earned citation is frame 4 for the measurer, and frame 3 only for the document that produced it: no delivery log, no denominator, nobody to sue.

So the prediction is uncomfortable and fairly safe. Within about two years the paid line on the answer surface will carry an audited, MRC-shaped number, and the earned line beside it a vendor index with a disclosure sheet. Both describe the same screen; only one is bankable in a budget meeting.

What that does to earned-media budgets

Not what most people assume. Earned citations dominate the evidence: Muck Rack’s May 2026 study of more than 25 million cited links across ChatGPT, Claude and Gemini found 84% came from earned media, and McKinsey’s work cited in the IAB framework puts a brand’s own website at only 5 to 10% of the sources AI platforms reference. The mechanism is not in question. The measurement asymmetry is, and budget follows the number that survives a finance review rather than the number that is true — which is how local citations came to be underfunded beside backlinks for a decade.

The one quantity in the room with a census frame

Go back to row three. A public record — published documents anyone can re-count — is a census frame. It is not a sample, so it carries no sampling error. It is dated, so precedence is checkable. It does not move when a vendor rebases, because no vendor issued it. And any stranger with a browser can verify it, which is what the whole standards apparatus exists to approximate.

For a business trying to be cited, that record has a name: the third-party documents stating the claim you want restated. Not a visibility score, not share of voice. The count of independent pages carrying the fact, and the date each appeared. It is the only number here that a court, an acquirer, a regulator and a competitor would all read the same way.

Why this is unusually convenient

In most measurement problems the auditable quantity and the causal quantity differ, and you end up managing the proxy. Here they coincide: the corroboration record is at once the most verifiable quantity available and, on the published evidence, the strongest input to the outcome. That is a fortunate accident worth exploiting, and it runs through everything from securing durable sponsorship placements to earning inclusion in third-party listicles and the editorial standards behind guest placements that survive review.

Three cautions. A census of documents is a count, not an effect: more of them proves nothing about whether an engine read any. Coverage error is real, since unlinked mentions and paywalled trade titles will be missed. And a document repeating what you published adds a row without adding a witness. None of that undermines the frame; all of it belongs in the footnote under the number.

THE CONTRACT CLAUSE

Four provisions for any agreement where money moves against an AI visibility number — or for an internal reporting policy.

  • Freeze the basket. The prompt list is an appendix, fixed for the term. Any addition creates a new series; both series are reported side by side for one full period before the old one retires.
  • Report a band, not a point. Establish the variability baseline before the term starts by re-running the frozen basket inside a seven-day window, and report every figure with that band attached.
  • Re-baselining suspends the KPI. A movement that appears across the whole category on one platform is the platform, not the supplier. Suspend, re-baseline, restate — never claw back.
  • Name one census-frame KPI. At least one target must be a countable, dated, third-party-verifiable record, so the contract does not rest entirely on an index whose owner can rewrite it.

What should an AI visibility KPI look like in a contract?

It should name the prompt list as an appendix, express the target as a range rather than a figure, state the engines and model versions, and pair the index with a countable third-party record. A KPI written as “visibility score to 30 by Q4” is a bet on a vendor’s release schedule.

Ashgrove Payroll Bureau: the year a KPI measured the instrument

Ashgrove Payroll Bureau in Chester runs payroll and CIS deductions for construction and facilities clients: 74 staff, £6.9M in annual fees. In January 2026 it signed a twelve-month agency retainer whose headline KPI was one line — raise the brand’s visibility score with a named tracking platform from 18 to 30 by the end of Q4, with 30% of the fee at risk.

By March the score read 24 and both sides were pleased. In April the vendor expanded its UK prompt library from 120 to 400 and added two engines to the default view. Ashgrove’s score fell to 15, below the starting baseline, with no change to its content, coverage or placements. The clawback triggered automatically and the finance director paused the programme for six weeks.

In May the agency re-ran the original 120 prompts itself, ten runs each: 1,200 calls at published list prices for £34 and three hours of an analyst’s time. On the original basket the brand scored 26. The entire recorded decline was the instrument. Nothing had happened to Ashgrove except that somebody else changed a list.

The rewrite

In June the contract was reissued against the four provisions above. A frozen basket of 140 prompts went into an appendix, covering all four intent types, with engines and model versions named. Results were reported as a band — 26 to 31, rather than 28 — computed from a variability baseline established by re-running the basket three times in the week before the term began. A second KPI was added that no vendor controlled: the count of independent third-party documents stating the two facts Ashgrove wanted restated, its CIS verification turnaround and its sub-contractor onboarding time. At signature that count was 41.

In July the vendor re-based again on a new model version and the headline score moved from 21 to 17. Nobody panicked: the frozen basket read 29 and the document count had reached 61, both traceable to a cause. By period end the census stood at 74 documents across trade press, two accounting bodies and a construction procurement register, 19 carrying the turnaround figure verbatim.

Four things that went wrong anyway

The frozen basket went stale: two prompts named a product Ashgrove retired in May, and change control kept them in until the period closed. The census undercounted, with at least eleven hand-found mentions unlinked and invisible to any backlink tool — the problem behind our work on references crawlers never see. Self-running the basket cost real time, from the analyst also writing the outreach. And the board kept asking for the vendor number anyway, because it produced a chart and a band does not.

Where this argument could be wrong

The strongest counter is not that vendors will improve. It is that the platforms hold the log and can hand over a census frame whenever they choose — and are already doing it. On 3 June 2026 Google added generative-AI performance reports to Search Console, breaking out AI Overviews and AI Mode impressions by page, country, device and date, to a subset of UK site owners first. Microsoft shipped an AI performance report in Bing Webmaster Tools in February 2026. If every platform issues a delivery census, the frame problem dissolves.

The counter gets stronger. The industry has transacted for a decade on Search Console impressions — unaudited, issued by the seller, accredited by nobody — and almost nobody complained. If that worked for search, an unaudited platform number may work here too, and the insistence on adjudicators looks fussy.

Where the counter stops

Three bounds, none fatal to the objection but all binding. First, the reports are impressions-only: no clicks, no click-through rate, no queries. They tell you that you appeared, not what appearing was worth — the term our analysis of what an agentic browsing click is worth showed to be unstable. Second, they cover your own pages inside one company’s surface, so a platform census can never produce share of voice: it cannot show where competitors appeared, and share of voice is what the category argues about. Third, a platform-issued census measures the platform’s product. Five platforms will issue five censuses on five definitions, and five numbers that do not reconcile are not a currency.

A second objection deserves an answer: standards do form without money — ISO, W3C, and content provenance work like C2PA. True, but those are interoperability standards, where the cost of divergence falls on implementers and adoption is self-enforcing. A measurement standard is nothing but adjudication. The provenance record makes the point: as we found examining Content Credentials for publishers and what C2PA can say about a link, a specification with no verifier produces enormous write volume and almost no reads — which is also why unsigned content is not the liability it is billed as.

A regulatory route could also shortcut this, and is worth watching rather than betting on. Disclosure duties under the EU’s AI Act content rules and the CMA’s UK remedies are already forcing platforms to expose data they would not have volunteered; the UK-first rollout of Google’s AI reporting was no coincidence. Regulation can manufacture a frame where commerce will not — on its own timetable, and granting nobody a right to a competitor’s number.

What to do on Monday

  • Write down your measurand. One sentence per number you report: which of the four quantities it is, on which engines, under which retrieval setting.
  • Ask every provider the disclosure questions — prompt library construction, query sourcing, model versions, retrieval on or off, re-baselining history. Treat a refusal as an answer.
  • Freeze your basket this week and put it in an appendix, even if the contract is signed. Nothing else this quarter has a better ratio of effort to defensibility.
  • Establish a variability baseline before interpreting any movement: re-run the frozen basket three times inside seven days and record the spread. That is your noise floor.
  • Strip point estimates out of board reporting and replace them with bands. Say the band aloud the first time so nobody reads it as imprecision.
  • Start the corroboration census now: every independent third-party document carrying the claims you want restated, with its date and publisher. It takes an afternoon and it back-dates.
  • Rewrite one KPI. Move at least one target off the vendor index and onto the census, so a rebase cannot cost anybody their fee.

The field asked for a benchmark, received a code of practice, and read the difference as a disappointment. It is not. A code of practice is what an industry gets before anyone is exposed enough to pay for an auditor, and the discipline it demands — say what you measured, how, and what moved rather than what was merely re-listed — is available to any team today. It is also the strongest position an earned-media practitioner has held in years, because the one input that survives every rebase is a document somebody else published with a date on it. The standard will arrive when someone can be sued over the number. Until then, count what a stranger can check.

If you are building the underlying programme rather than the reporting layer, start with the fundamentals of link building, work through the current strategy set, and keep the 2026 statistics baseline beside you when someone quotes a benchmark in a meeting. If you are hiring for it, the specialist brief has changed accordingly.

Leave a Reply

Your email address will not be published. Required fields are marked *

Experience-Led Content Engine Previous post Building an Experience-Led Content Engine That AI Rewards
AI Visibility Distribution Next post Visibility Is a Distribution, Not a Number: Sampling AI Answers Properly