TL;DR
An audit was never a measurement exercise. It was fault localisation — naming which stage failed, who owns it, and what the repair costs. A results page gave every absence an address. An answer engine gives none, so the audit sold across 2026 observes a terminal outcome and then prescribes to a stage it never tested.
Britain now has the best first-party AI reporting in the world, and it still cannot tell you why you are missing, because absence leaves no first-party trace. This piece rebuilds the audit as an elimination procedure: six places an absence can live, cheap probes that tell them apart, and a closure rule that survives a noisy re-test. Three of the six sit on somebody else’s property — which is why the deliverable stops being a development backlog and becomes an acquisition brief.
The best AI reporting in the world still cannot tell you why you are missing
On 3 June 2026 the Competition and Markets Authority (CMA), Britain’s competition regulator, imposed a publisher conduct requirement on Google under the Digital Markets, Competition and Consumers Act. One obligation reads like a gift to anyone running an audit: Google must give publishers clear and detailed metrics on how users engage with their content inside search generative AI features.
The same day, Google launched Search Generative AI performance reports in Search Console — rolled out to a subset of UK site owners first. Impressions inside AI Overviews and AI Mode, separated from ordinary web results for the first time, sliceable by page, country, device and date, with data beginning 18 May 2026 and no historical backfill. British site owners therefore hold something no other market has: a regulator-secured first-party view of a surface everyone else is guessing at, on a search engine carrying 91.47% of UK queries (Statcounter, April 2026).
Then read what shipped. Impressions. No clicks, no click-through rate, no average position, no queries. Google’s own announcement notes that this data was always inside your overall performance totals; the new view separates it rather than adding to it.
The reflex reading is that the report is unfinished and the answer arrives when clicks do. That assumes the missing columns are the problem. They are not. A completed version of this report would still not answer the question every board actually asks, which is never how often did we appear but why were we not there.
It is worth sitting with the gap between the obligation and the artefact. The regulator asked for engagement metrics and the first delivery was an appearance count, with the substantive publisher controls phased across a calendar running into 2027. That is not bad faith; it reflects how hard the underlying reporting problem is on a generated surface. But it means the most instrumented AI search market on earth has, as its flagship instrument, a report that can confirm you were present and can say nothing at all about the far larger set of answers in which you were not.
Absence leaves no first-party trace
A Search Console report is a report about your own pages. It can only contain rows for pages that appeared. The page that was never fetched produces no row. The competitor comparison table that answered the question instead produces no row. The model that answered from what it already believed, without searching at all, produces nothing anywhere in your estate — not in your logs, not in your analytics, not in the regulator’s mandated metrics. You cannot audit an absence from inside the thing that is absent.
What is a SERP-less audit? A SERP-less audit is the attempt to establish where a brand stands in AI-generated answers when there is no ranked, persistent, shared results page to inspect. Instead of reading a public artefact, the auditor generates answers by prompting, records what appears, and infers what to change. The recording is easy and now heavily tooled. The inference is the hard part, and it is the part the field has quietly skipped.
What an SEO audit was actually doing all along
Before rebuilding the thing, be precise about what it did. The classic technical SEO audit felt like measurement — scores, percentages, red and amber rows — but measurement was the packaging. The work was fault localisation.
A page’s journey to a click passed four observable checkpoints, and each had its own independent instrument. Fetched, visible in server logs. Indexed, visible in Index Coverage. Ranked, visible in a rank tracker. Clicked, visible in Search Console. If you were absent from the results page, one query against those four instruments told you which checkpoint had failed. Crawled but not indexed is a quality or duplication problem. Indexed but not ranking is a competitive or link authority problem. Ranking but not clicked is a title and snippet problem. Every absence had an address.
Why a classic finding was repairable
Four things were true of a classic audit finding at once. It was deterministic: a 404 is a 404 on every request. It was attributable: it belonged to a named stage. It was verifiable: you fixed it, re-tested, and the defect was gone permanently. And it was internal: nobody’s permission was required. Even the harshest sanction in the discipline behaved this way — a manual action arrived with a named cause, an affected scope and a documented route back, and the disavow route was a lever you could pull unilaterally.
That combination is why nobody ever bought a site health percentage. They bought the list of located causes and the order to fix them in. The score was a wrapper for a repair schedule.
The word audit carries that meaning everywhere else it is used. A financial audit does not score a company out of a hundred; it locates discrepancies and names them. A safety audit does not report a percentage; it identifies specific failure modes and who must close them. Only in search marketing has the word drifted toward a rating.
Key takeaway
An audit is a diagnostic instrument, not a thermometer. Its output is a located cause with an owner and a cost, not a number describing the patient. Any audit that ends in a percentage plus a generic improvement list has skipped the only step that made auditing worth paying for.
The playbook everyone is running, stated at full strength
The 2026 AI visibility audit has settled into a stable shape, and it is worth stating generously before taking it apart. Build a prompt set covering branded, category and commercial-intent questions. Run it across ChatGPT, Gemini, Google AI Mode, Perplexity and Copilot. Record whether the brand is mentioned, in what position, with what framing, and which sources are cited. Run the same set for two or three competitors and compute a share-of-voice comparison. Check accuracy and sentiment where you do appear. Confirm the site is machine-readable. Then deliver recommendations: structured data, question-shaped headings, author bylines with credentials, consistent brand descriptions across profiles and listings, presence on Wikipedia and Wikidata, and more third-party mentions.
Much of this is honest, useful work. It establishes a baseline where none existed. It is candid about the absence of a Search Console for ChatGPT — one widely circulated agency methodology opens by saying there is no such console, no standardised metrics and no consensus on what to look at, which is a more truthful preamble than most of the industry manages. And it correctly identifies the commercially interesting cell: the prompts where a competitor appears and you do not.
The join that is never made
The break comes at the handover from observation to prescription. The audit observes a terminal output — you were absent, they were present — and then prescribes to a stage of a pipeline it never inspected. There is no observation in between that discriminates between the possible causes.
Read the reasoning as it is actually published. A representative 2026 audit guide advises that if a competitor’s answers appear cleanly and yours do not, the problem is usually in the content. That word is doing enormous load-bearing work. It is not a finding. It is a prior — and, more often than anyone admits, it is a prior shaped by what the auditor sells. A content agency finds a content gap. A technical shop finds a schema gap. A digital PR firm finds a mentions gap. All three ran the same probe and got the same null.
A presence score tells you that you are absent. It does not tell you where you are absent — and every remedy on the market is a remedy for a different where.
The Absence Differential: six places a missing brand can be missing from
Clinicians do not treat a symptom by listing every disease that produces it. They order the cheapest test that rules the largest number out. That is the move the SERP-less audit needs, because a single null result — brand not mentioned — is compatible with six structurally different failures, each with a different owner, a different cost and a different clock.
The table below is the instrument. Read the second column first, and notice how little separates the rows: in a standard audit run, all six look nearly identical. The third column is where the work happens. Each probe is a few prompts and a few minutes, and each one eliminates at least one row.
| Where the absence lives | What the standard audit records | The discriminating probe | Owner and clock |
| 1. Not retrieved | Brand absent; competitors named; no citation to your domain. | Did the engine search at all, and did any of your URLs enter the cited set for any prompt in the run? | Discovery and corroboration. External. One to two quarters. |
| 2. Not legible | Brand absent; competitors named; no citation to your domain. | Fetch the page the way a retrieval client does — no JavaScript execution, no cookies. Is the claim still in the text? | Engineering. Internal. Days. |
| 3. Not corroborated | Brand absent, or named without a link to you. | Name your brand in the prompt. If it can answer accurately only when told, you are legible but not selected. | Earned coverage. External, with consent. One to two quarters. |
| 4. Wrong shape | Present for definitional prompts, absent for comparison, eligibility and shortlist prompts. | Split the prompt set by answer shape and compare presence rates across the two halves. | Content and third-party listings. Shared. Weeks to a quarter. |
| 5. Answered from weights | Brand absent and no sources cited at all, or only long-established ones. | Re-run with retrieval explicitly forced. If you appear only then, the unforced answer never consulted the web. | Nobody, directly. Long clock, indirect. |
| 6. Used, not named | Your figures, phrasing or framing appear in the answer with no attribution to you. | Search the answer text for claims only your estate publishes, then check the citation list. | Attribution and framing. Shared. Weeks. |
Row six deserves more attention than it gets, because it is the only null that is not a null. Your claim is in the answer and your name is not, which means the estate is doing retrieval work and losing the attribution. That is a framing and citability problem rather than a coverage one, and it is often the cheapest win available: a distinctive owned asset that is hard to paraphrase without naming its source converts silent use into a citation more reliably than another explainer page does.
Run the six in order and the search space collapses quickly. Rows two and five are eliminated or confirmed in under an hour, at no cost. Row one usually needs a second look at whether you are in the candidate corpus at all — which is a question about what AI systems treat as a source, not about your copy. Row three is the expensive one, and it is also the most common, which is exactly why an audit that skips the probe and prescribes from a prior tends to prescribe the cheap rows.
The forced-mention probe: separating never seen from seen and passed over
Of the six discriminations, one carries more weight than the rest, because it splits the two failures that look most alike and cost most differently. Ask the same commercial question twice: once as a buyer would, and once with your brand named in the prompt.
The forced-mention probe
Run A — the buyer’s question, brand not mentioned. Record presence and citations.
Run B — the same question with your brand named. Record what the engine can say about you, and where it sourced it.
Then read the pair:
Accurate, specific and current in B, absent in A — you are retrievable and legible. The failure is selection. This is a corroboration problem, not a content problem.
Vague, hedged or refused in B — you are not reachable or not readable. Fix that before spending a penny on coverage.
Confident but stale or wrong in B — an in-weights prior is answering and retrieval is not correcting it. Freshness and volume of recent third-party mentions, not on-page edits.
Accurate in B but sourced entirely to third parties, none of them you — the engine knows exactly who you are, from documents you do not control. Your own estate is not in the candidate set even when you are the subject of the question.
That fourth outcome is the one nobody looks for, and in practice it is common for established firms. It also inverts the usual reading of a good result: an answer that describes you accurately is not evidence that your content is working. It may be evidence that other people’s pages are working, and that yours are decorative.
One methodological rule makes or breaks the probe: run A and B in the same session state, the same account tier and the same region, within the same short window. A comparison between a logged-out run on Monday and a signed-in run on Thursday is not a comparison. Micro-example: a Bristol accountancy firm ran A on a colleague’s personal account and B on a clean browser, concluded it had a content problem, and spent eleven weeks rewriting service pages. The difference was the account.
Key takeaway
Never seen and seen-then-rejected produce the same number in every commercial tool on the market, and they have opposite remedies. One is an engineering ticket or a discovery problem; the other is a quarter of earned coverage with somebody else’s editorial consent attached. Two prompts tell them apart. No dashboard does.
What no probe can establish, and why the audit must stop claiming it
Elimination narrows the cause. It does not make the reading representative, and honesty about the gap is what separates a diagnostic from a sales document.
The audit measures a surface no returning customer uses
The field’s own best practice is to run prompts in logged-out or temporary chats for cleaner, more repeatable results. That advice is correct on its own terms and quietly damning: it selects a surface for reproducibility, not for representativeness. The reproducible number is the cold one — how a model describes you to a stranger with no history.
The warm surface has been moving fast in the other direction. OpenAI’s May 2026 memory update widened what can be pulled into personalised context; ChatGPT’s Fast Answers, live since April 2026 across logged-in and logged-out users, added a second answer pathway on factual prompts so the same question can now resolve two different ways; and in July 2026 the custom-instructions limit rose to 5,000 characters, making standing personalisation a persistent brief rather than a preference. One tracking vendor has started splitting this into a cold visibility problem and a personalised one. The audit you can reproduce is a measurement of the first.
Your probe account has no age, no history and no postcode
Ofcom’s Adults’ Media Use tracker put UK generative-AI use at 54% in late 2025 against 31% a year earlier, but with a spread that should stop any auditor treating a single account as a population: 79% of 16-to-24s and 74% of 25-to-34s, against 18% of over-65s. Ofcom’s Online Nation 2025 recorded 1.8 billion UK visits to ChatGPT in the first eight months of 2025, up from 368 million, and roughly 30% of UK keyword searches returning an AI-supported response. Region matters as much as age — engines use different indexes and different availability by market, which is why audiences outside your home market routinely see a different source set for the same words, and why a single global visibility figure is a category error.
Can I audit AI presence from Search Console alone? No. Search Console can tell you where you appeared inside Google’s own AI features, which is a genuine and newly available signal, but it is silent on every other engine and structurally silent on absence. It is a confirmation instrument, not a diagnostic one — useful for verifying that a fix landed, useless for deciding which fix to make.
Close findings on inputs, not on outputs
This is the operating rule that follows. A classic finding closed on a re-test: fix, re-crawl, confirm, done. A SERP-less finding cannot close that way, because the re-test moves on its own. Re-running an unchanged prompt set weeks later can shift presence several points in either direction with no work in between.
So close on the input. The finding closes when the corroborating document is published, the register entry is live, the page renders server-side, the dated study exists at a URL. Whether presence moved is a separate, slower and statistical question, and it belongs in a different report on a different cadence. Conflating them is how teams end up chasing a citation loss that was noise, and how good remediation gets cancelled in month two.
Why the findings keep landing on somebody else’s property
Run the differential across a real prompt set and the loci do not distribute evenly. Rows two, four and six are wholly or partly yours — rendering, page shape, framing. Rows one, three and five are not. And the two groups differ in more than ownership.
The cheap loci are gates; the expensive one compounds
Legibility is a threshold. Server-render the claim, put the fact in text rather than an image, keep the rendering path clean — and then you are done, permanently and at parity with everyone else who has also done it. There is no advantage in a gate that every competent competitor clears. It is worth clearing precisely because it is cheap, and worth nothing more.
Corroboration behaves the opposite way. It is slow, it is bought with editorial consent rather than money alone, it has a refusal rate, and it accumulates. That asymmetry is why the differential matters commercially: if you cannot tell the gate failure from the corroboration failure, you will keep spending on the gate, because the gate is the part you can finish.
The deliverable changes genre
An audit whose findings sit off-property is no longer a development backlog. It is an acquisition brief. A 404 needs nobody’s permission; a corroboration finding needs a third party to agree, and that changes the document you hand over. Each finding now carries a target class, a lead time, a refusal allowance and a sensible acquisition pace — closer to a media plan than a ticket queue. Targets stop being chosen by domain metrics and start being chosen by which locus they close: a small community or practitioner surface that answers a selection question can outperform a far stronger domain repeating an explainer, because only one of them vacates a row in the differential. It also breaks the standard calendar. Quarterly re-audits exist because dev tickets clear in a quarter; off-property remediation runs on eight-to-twenty-week clocks with refusals, so a quarterly re-audit routinely measures work that has not landed and then reports it as a failure of the work.
Why does my brand appear in Google AI Overviews but not in ChatGPT? Because they are different systems with different corpora, different retrieval rates and different priors, so the same absence has different causes on each. Before concluding it is a content problem, establish whether the engine searched the web at all on that prompt — an engine answering from weights is telling you about its training data, not about your site.
Which is also the honest limit on where any of this can be fixed. Row five has no owner you can email. The only lever on a parametric prior is a long, slow accumulation of the kind of third-party record that survives into the next corpus — the same slow substrate that decides which sources surface in AI Overviews, measured in years rather than campaigns.
Worked example: Ashcombe Risk Partners
Ashcombe Risk Partners is an invented composite, built from the pattern this differential keeps surfacing. A Manchester commercial insurance broker specialising in haulage, cold-chain and freight-forwarding risks: £38M gross written premium placed, £5.2M income, 31 staff, and roughly a third of new business historically originating in search.
Their February 2026 audit, bought from a capable agency, ran 50 prompts across four engines and returned a 16% presence rate against a market leader’s 44%. The recommendation set was the standard one: schema, FAQ blocks, author credentials, a Wikidata entry, and a content programme targeting the prompts where they were absent. They spent £31,000 over four months. Presence in June: 18%, inside the noise band.
What the differential found
Re-run as an elimination procedure in July, on the same 50 prompts, the nulls sorted very differently.
- Row two, legibility: clean. Their key pages server-rendered, claims in text. Eliminated in 40 minutes.
- Row five, weights: 9 of 50 prompts returned no web search on at least one engine. Those nulls carried no information about the site at all and were removed from the denominator, which alone moved the honest baseline from 16% to 19.5%.
- Row three, corroboration: the forced-mention probe returned accurate, current, specific answers on 4 of 4 engines — sourced to a trade title, the British Insurance Brokers’ Association member directory, the Financial Services Register and two client press releases. Not one citation to ashcomberisk.co.uk.
- Row four, shape: presence was 38% on definitional prompts and 4% on the broker-selection and eligibility prompts that actually precede an enquiry.
The fourth outcome of the forced-mention probe was the finding of the engagement. Engines knew Ashcombe well and had never once used an Ashcombe page to say so. Four months of content spend had been aimed at a locus that was not failing.
Remediation, and what it cost
Over 19 weeks: a dated claims-handling-time study drawn from their own 2023 to 2025 case records, published with methodology; two sector register entries; a cold-chain risk guide co-published with a logistics trade body; and 11 placements commissioned against the eligibility and selection prompts specifically, rather than against domain-authority targets — two buying-group listings, three trade titles, a regional business title, two freight association pages and three broker-comparison resources.
At 19 weeks: presence on the corrected denominator 19.5% to 41%, selection-prompt presence 4% to 27%, own-domain citations 0 to 14 across the run, and inbound enquiries citing an AI recommendation 3 to 11 per month.
The negatives were real. The differential itself took 26 analyst-hours, so it is a quarterly instrument at best, not a monthly one. Two of the 11 commissioned placements were refused outright and a third took five months. The claims-handling study drew a compliance review, because published performance figures from an FCA-regulated firm carry promotion rules that a marketing team is not qualified to sign off — two claims came out of the final version. And one probe read wrong: a null they diagnosed as corroboration turned out, six weeks later, to be a legibility failure on a subdomain nobody had included in the fetch test. The elimination procedure is only as good as the estate you point it at.
The strongest objection: just do everything
The hardest counter to this entire argument is not that the six loci are wrong. It is that localising them is a waste of time. It runs like this, and it deserves its full strength:
Differential diagnosis is what you do when treatment is expensive, risky or mutually exclusive. None of that applies here. Structured data is a day. Server-rendering a template is a sprint. Content is being produced anyway. Earned coverage is something a competent firm should be doing regardless. The remedies are cheap, additive and non-exclusive, so shotgun therapy dominates: do all six, skip the diagnostics, and you will have fixed the real cause plus five harmless others in the time the elimination procedure takes to design.
That is how most competent teams actually behave, and on a two-locus problem they would be right. Four things bound it.
First, the remedies are not comparably cheap, and the cost distribution is severely skewed. Hours for legibility; a quarter, a budget and another organisation’s consent for corroboration. In practice, do-everything resolves to do-the-cheap-things-and-defer-the-expensive-one, because the cheap things finish and the expensive one does not. Shotgun therapy under a skewed cost distribution is not shotgun therapy. It is selection on cost, and it systematically under-fires on the most common locus.
Second, simultaneous intervention destroys the learning. With a noisy outcome and six changes shipped together, nothing is attributable, so next quarter you re-run the same audit with the same non-information and make the same guess again. Ashcombe’s first £31,000 bought no knowledge, which was the more expensive loss.
Third, the remedies are not harmless. Publishing performance figures invites regulatory review in regulated sectors and hands rivals a benchmark. Consolidating thin pages sheds long-tail rankings. Chasing volume on the corroboration row without a target class buys placements that occupy no locus at all, which is the most expensive way to be busy. Internal-only fixes have their own version of this: a decade of equity-shuffling on-site is what the reflex to solve everything on your own property looks like when it runs unchecked.
Fourth, and decisively: the shotgun is only rational when there is no cheap discriminator. Here the diagnosis costs less than the cheapest treatment. Two prompts and forty minutes beat a day of schema work on price. When the test is cheaper than the therapy, empiricism wins on the objection’s own economics.
None of which makes the objection foolish. On a site with an obvious legibility defect, diagnose-then-treat and treat-everything converge, and the faster path wins. The argument here is narrower and holds where it matters: the moment one candidate remedy costs an order of magnitude more than the others and depends on somebody else saying yes, guessing stops being efficient.
What would falsify the argument: a controlled run in which brands matched on legibility and page shape, differing only in third-party corroboration, show no difference in unprompted presence. Any team with two comparable client sites can run it in six weeks, and it is the test this piece should be judged on.
What to do on Monday
A literal sequence. The first four cost nothing but time and will eliminate half the possibilities before you commission anything.
- Pull your prompt set apart by answer shape — definitional, comparison, eligibility, shortlist — and record presence separately for each. One blended number hides the split that matters.
- For every null, record whether the engine searched the web at all. Remove the no-search nulls from the denominator and report that count openly as a separate line.
- Fetch your five most commercially important pages without JavaScript, cookies or a logged-in session, and confirm the load-bearing claim is still present as text.
- Run the forced-mention probe on ten commercial prompts, in matched session state, and classify each pair against the four readings.
- Count how many of the citations in your accurate branded answers point at documents you do not own. That percentage is your real dependency figure.
- Rewrite every finding to name its locus, its owner and its clock. Any finding you cannot assign an owner to is not yet a finding.
- Convert the off-property findings into a commissioning brief with target classes, lead times and an explicit refusal allowance — and stop reporting them on the same cadence as engineering tickets.
- Set closure criteria on inputs. Presence re-measurement moves to its own quarterly cycle, with a stated noise band, and is never used to approve or cancel work inside it.
The results page did the diagnostic work for the discipline for twenty-five years, and its disappearance took the diagnosis with it, not merely the number. Building a modern link and citation programme on a presence percentage is building on the one output that cannot distinguish a broken template from an absent reputation. The remedy is not a better dashboard. It is the older discipline the dashboards replaced: rule things out, cheaply, in order, before you spend.
If the differential lands on the corroboration row, the working definitions are in the complete beginner’s guide to link building; the instrumentation options are in the tools review; and the base rates worth arguing from are in the 2026 statistics roundup.
