Position to Probability

From Position to Probability: What “Ranking” Means in 2027

TL;DR

A rank had five properties, and only one of them — ordinality — was ever about performance. The other four were what made the number admissible: universally observable, attached to a query, attached to a URL, and reproducible.

Those four failed one at a time over fifteen years. What finished in 2026 was not measurement accuracy. It was intersubjectivity — the property that let a client, an agency and a competitor read the same number and agree it said something.

The replacements do not agree with each other either. BrightEdge found ChatGPT, AI Overviews and AI Mode disagree on brand recommendations for 61.9% of queries; Superlines measured citation-rate variance across engines reaching 615x.

So this is a contracting problem before it is a measurement problem. SEO was priced for two decades against a number supplied by a third party both sides trusted. That oracle has stopped publishing.

Instruments below: the Five Properties of a Rank (what each one bought you, what broke it, and what replaced it) and the Three Witnesses (who can actually verify a number — and therefore what you can put in a performance clause).

1. What a rank actually was

For twenty years the search industry treated a ranking as a performance measurement. It was really something stranger and more useful: a publicly issued fact about a private system, produced by a party with no stake in the argument you were having about it.

That is an unusual thing for any profession to have. Very few marketing disciplines are graded by an oracle that both the buyer and the seller of the work consider neutral. Rank tracking outlived twenty years of people calling it a vanity metric precisely because performance was never its main job. Its main job was settlement: deciding who was right when a client and an agency disagreed about whether the work had worked.

What is a ranking, structurally? A ranking is an ordinal position in a published list, produced deterministically from a stated query, attached to a specific URL, and observable by anyone who runs the query. Five properties. Strip any one of them out and the number still exists, but what you can do with it changes completely.

And of the five, only ordinality — the fact that you were above or below somebody — measured anything about your performance. The other four measured nothing at all. They were infrastructure. They were what made the number admissible as evidence.

It is worth being concrete about what that bought. A rank-based retainer, a page-one guarantee, an in-house budget defence, an agency pitch, a scorecard in a quarterly review, a dispute over whether an eighteen-month programme had delivered — all of these ran on a figure that neither party produced and neither could edit. Strip that out and the commercial machinery of an entire profession loses its arbitration mechanism, which is a considerably bigger problem than losing a chart. It is also why the role of the link building specialist is being rewritten from the contract inwards rather than from the tactics outwards.

Is keyword ranking still worth tracking in 2027? Yes, but for diagnosis rather than for settlement. A position still tells you whether a page is retrievable, indexed and topically resolved; it no longer tells you what share of the answer you hold, and it should not appear in anything you sign.

2. The five properties failed one at a time

This is the part that gets lost in the 2026 conversation, which tends to treat the collapse as a single event dated to an AI launch. It was not. Four of the five properties had already gone, and the industry had built a workaround for each one and then forgotten it was a workaround.

THE FIVE PROPERTIES OF A RANK

What each property bought you, what removed it, and what the industry substituted.

PropertyWhat it bought youWhat broke itWhenThe substitute
UniversalityOne list everyone could see — client, agency and rival read the same numberPersonalised results, then localised ones2009–2012A simulated “neutral” location and logged-out state that no real user occupies
Query attachmentA traceable line from a phrase people typed to the work you didSecure search and “(not provided)”2011–2013Sampled query data; Ahrefs measured 46.77% of clicks anonymised
ReproducibilityRun it again, get the same answer — so a dispute could be re-runNon-deterministic generation2024–2026Repeated sampling and averaged rates, per SparkToro’s guidance
URL attachmentAccountability to one asset you owned and could fixQuery fan-out and passage-level retrieval2025–2026A claim inside an answer, assembled across domains
OrdinalityThe only property that measured performanceAI Overviews, then AI Mode as default2024–2026Citation presence — binary per run, not a ladder

Read the substitute column downwards and the pattern is obvious in hindsight. Every time a property failed, the industry replaced a fact with a model — a simulated location, an estimated query set, an averaged rate — and carried on reporting the output as though it were still the original kind of object. Rank tracking survived by becoming a simulation of a thing that had stopped existing, and it worked for a decade because everyone ran the same simulation.

The ordinality row deserves its own numbers, because it is the row people dispute. Ahrefs’ December 2025 study of 300,000 keywords against Search Console data found position-one click-through rate at roughly 1.6% where an AI Overview appeared, against 7.3% two years earlier — a 58% fall, up from 34.5% measured in April 2025. The detail that matters more: position-one CTR had also fallen to 3.9% on queries with no AI Overview at all, a 49% decline. Position one is not one thing that got worse. It is several different things wearing the same label, and the label stopped distinguishing between them. That is the practical content of what AI Overviews did to the value of a ranked link.

Key takeaway

The 2026 break is not that ranking got harder to measure. It is that the last property with any performance content — ordinality — went at the same moment as the last shred of shared observability. Before, you had a flawed number everyone read the same way. Now you have several sophisticated numbers that nobody reads the same way.

There is a second reading of that column worth naming, because it explains why the field keeps being surprised. Each substitution was individually reasonable and collectively catastrophic. Nobody made a bad call when they adopted a neutral tracking location, or accepted sampled query data, or started averaging repeated runs. Each was the best available response to a specific loss. The cost only appeared in aggregate, when the fifth property went and someone finally asked what the surviving number was actually made of.

3. What broke was agreement, not accuracy

It is tempting to frame the loss as precision. It is not. Citation-share estimates are in some respects better evidence than a rank ever was: they describe what a real answer contained rather than what position a link occupied on a page most people never scrolled.

The thing that has actually gone is duller and far more consequential. A rank was intersubjective. Two parties with opposed interests could observe it independently, at the same moment, from a source neither of them controlled, and arrive at the same value. That is a rare property. It is the property that lets a number appear in a contract.

Citation probability has none of it. Each observer samples the system separately, at different times, with different prompts, through different panels, and gets a different answer — legitimately, because the system is non-deterministic. Nobody is wrong. There is simply no shared reading to be had. Brainlabs put the point precisely in April 2026: the deeper problem is not the tools, it is the medium.

For example: an agency reports 34% citation share; the client’s in-house analyst runs the same prompts a fortnight later and gets 19%. Under a rank regime that gap was a bug, and somebody was misconfigured. Under this regime it is the expected behaviour of the measurement, and there is no third party either side can appeal to.

It is worth resisting the instinct to call this a defect that will be engineered away. Non-determinism is not a rough edge on generative retrieval; it is a consequence of how these systems sample, and the same property that makes an answer feel responsive is the property that makes it irreproducible. Ekamoira found roughly 73% of fan-out queries change between runs of an identical prompt. SparkToro’s guidance — repeat the runs and average the rates — is the correct methodological response and it is also an admission: the underlying event you want to observe does not have a single value waiting to be read.

Which is why the disappearance of a citation is such an awkward thing to investigate. Under a rank regime, a drop had a cause you could hunt. Now a page can be absent from three consecutive runs and present on the fourth without anything having changed on either side, and diagnosing genuine citation loss means first establishing that there was a loss rather than a sample.

4. The replacements do not agree with each other either

This would matter less if the successor metrics were converging on a standard. They are not, and the divergence is structural rather than early-days sloppiness.

BrightEdge’s cross-platform work found ChatGPT, Google AI Overviews and Google AI Mode disagree on brand recommendations for 61.9% of queries. Foglift found 61.7% of top-25 cited domains appear in exactly one engine’s top-25 list, with ChatGPT and Perplexity citation rates correlating at r=0.78 while both correlate with AI Overviews at only r=0.54. Superlines measured citation-rate variance across engines reaching 615x — Grok citing sources on 27.01% of responses, Perplexity 13.05%, ChatGPT 0.59%, Claude 0%.

The engines are not one process with different settings. They are different processes, and a blended “AI visibility score” averages across them as though they were comparable. Kevin Indig’s Growth Memo halftime report in July 2026 recorded ChatGPT’s share of AI-search activity falling from 78% to 56% in six months while Gemini climbed to 30% — a shift a single blended figure has no way of showing you. Profound’s July 2026 analysis found Claude invoked a live web search for only 36.6% of tested prompts, so for most of the time you are “tracked” against Claude you are not being tracked against a search engine at all.

None of which means the panels are worthless. Semrush’s expanded 2026 AI Visibility Index analysed 126 million prompts and produced genuinely useful structural findings — that brands holding stable scores did so on the back of consistent descriptions across independent third-party sources, and that organisations integrating AI visibility with existing SEO and communications work reported gains at 81% against 36% for those running them separately. Muck Rack’s May 2026 analysis of more than 25 million cited links across ChatGPT, Claude and Gemini found earned media accounted for 84% of all AI citations. Those are directional facts about how the world works. They are not measurements of your account, and they cannot arbitrate a disagreement about it.

The benchmark studies do not share a unit

Three cross-industry benchmarks published in 2026 illustrate the problem by existing. Pondral scored 200 brands across five verticals and 8,215 results, reporting a mean of 55.8 out of 100. Foglift’s Q2 2026 benchmark evaluated 4,217 brands with 150+ industry prompts and reported a SaaS median of 62 out of 100. Presenc monitored 2,847 brands continuously across five engines for at least ninety days. Different methodologies, different scoring models, different panels — and no way to convert one into another.

Practitioners have noticed. Digiday reported in May 2026 that marketers were questioning the cost of AI visibility platforms as inconsistent results fuelled scepticism, with one agency testing around seventeen tools simultaneously and citing a lack of benchmarks, inconsistent results and attribution difficulty. Spending continued anyway, which tells you the demand is real. It is a demand for an oracle, and what the market is selling is panels. Those are not the same product, and the tooling landscape will not resolve the difference by adding features.

5. The three witnesses

Once you see the problem as evidential rather than statistical, the practical question changes. It is no longer “how accurate is this number?” It is “who can check it?” — because that determines what the number is allowed to do.

THE THREE WITNESSES

First witness — the public artifact. Both parties can verify it independently, from a source neither controls. A live URL carrying your placement. A trade register entry. A published dataset with a date on it. Test: could a stranger reproduce it in a browser in sixty seconds? Contract on these.

Second witness — the disclosed figure. One party holds it and shows the other. Search Console, analytics, CRM. Verifiable in principle, but only by whoever owns the login, so it settles nothing in a dispute unless the method is agreed in writing first. Report these, with the definitions, date ranges and filters written down — that is what promotes a disclosed figure towards a public artifact.

Third witness — the attestation. Neither party can check it; you are both trusting a vendor’s panel. AI visibility scores, prompt volumes, share-of-voice indices. The tell: ask the vendor to publish panel composition, run frequency and prompt selection. If they will not, you have an opinion with a decimal point. Use these for direction. Never settle on them.

The rule: contract on first-witness facts, report second-witness figures against a written method, and keep third-witness numbers out of every performance clause you sign.

The complication, stated up front: first-witness status decays. Publications rewrite pages, CMS migrations drop links, registers get restructured. A public artifact needs a re-check cadence, not a one-time check.

Promoting a second witness

The middle tier is where most of the available improvement sits, and the work is unglamorous. A disclosed figure becomes near-first-witness when the method is specific enough that the other party could execute it themselves and land within a tolerance you both agreed in advance. That means naming the property, the date range, the filters, the exclusions, who runs it, on what day, and what counts as a miss. Publish the misses — a sample that only reports appearances is a marketing asset, not evidence.

For example: “citation share was 31%” is a third witness if a vendor produced it and a second witness if you did. It approaches a first witness only when it reads “60 named prompts, five runs each, three engines, run by the client analyst on the first Monday of the quarter, 93 appearances from 300 observations, full log attached”. Nothing about that is sophisticated. It is simply written down, which is the entire difference.

6. What still counts as a first-class witness

Run your current reporting pack through that sort and the result is uncomfortable. Most of what fills an AI-era marketing dashboard is a third witness. Almost everything else is a second witness with an unwritten method.

One category comes through intact, and it is not the one the discipline has spent two years talking about. An earned placement is a public artifact. It is a statement published by an independent party, at an address anybody can load, carrying a date, capable of being archived. Your client can check it. Your competitor can check it. A tribunal could check it. It does not require you to trust a panel, and it does not evaporate when a model is retrained.

That is not an argument that links matter more than they did. It is a narrower and more specific claim: of the assets an AI-era programme produces, third-party links and the placements that carry them are close to the only ones that survive the loss of a shared oracle, because they were never issued by the oracle in the first place. A citation is an event inside a private non-deterministic system, observable only by sampling. A placement is a document. The same asymmetry is why a competitor’s backlink profile remains one of the last genuinely checkable competitive facts in the discipline — you can audit theirs, and they can audit yours.

Not every placement is equally durable as an artifact, and it is worth sorting them on that basis rather than on estimated authority. Register entries, accreditation listings and trade-body membership pages are the most stable: they are maintained as records rather than as content, and nobody rewrites them for engagement. Editorial coverage in established publications is next. Inclusion in ranked roundups and listicles sits lower, because those pages are refreshed aggressively and your position within them is somebody else’s editorial decision on a rolling basis. Assets you host yourself — original calculators and tools that attract references — are durable on your side but depend on remaining fetchable, which makes the technical layer underneath link acquisition part of your evidence chain rather than a separate workstream.

There is a quieter consequence for how programmes get planned. If the only artifacts that survive scrutiny are produced by other people publishing about you, then the acquisition of those artifacts stops being a downstream tactic funded out of whatever content leaves over, and becomes the part of the programme that generates your evidence. That is an argument about reporting rather than about rankings, and it arrives at the same place from a direction the discipline has not usually taken.

The obvious objection, taken seriously

If you contract on links, you have created a reason to buy links. That objection is correct and it is the strongest thing anyone can say against this section. A publication that sells the KPI will find sellers, which is the mechanism that produced fifteen years of penalties and the entire literature on recovering from a manual action and when disavowal is actually warranted.

The corruption comes from counting, not from verifiability. A clause that pays per referring domain is a purchase order. A clause that specifies named publications from an agreed tier, editorially independent, dated, retrievable, and reviewed at renewal is a specification. Both are first-witness facts; only one creates a market for junk. If your contract cannot tell those apart, the problem is the clause, not the asset — and monitoring the rate at which new links arrive is a hygiene check on your own incentives as much as on anyone else’s.

7. The regulator is trying to rebuild the oracle

There is one serious attempt underway to restore a shared reading, and it is not coming from a vendor. The Competition and Markets Authority imposed a conduct requirement on Google in June 2026 — the first binding one of its kind anywhere — following Google’s designation with strategic market status in general search.

Read what it actually asks for. Publisher controls over AI grounding at domain and page level, phased through December 2026 and March 2027. Proper attribution with clear links. And “clear and detailed metrics on user engagement” in generative AI features. It does not ask for query data, and it does not ask for a ranking.

That is a revealing choice. The world’s first binding remedy in this space treats the missing thing as credit, not knowledge — and credit, delivered as an attributed link, is a first-witness artifact by construction. If the remedy lands as drafted, the object it makes newly verifiable is the citation-with-a-link, not the citation-as-a-score. Similar attribution and transparency logic runs through the EU’s content and AI transparency regime, which is worth tracking for the same reason.

Google’s own move in the same direction is instructive for what it withholds. The generative-AI performance reports launched in June 2026 — in a UK-first beta, following the regulatory pressure — report impressions, pages, countries, devices and dates. They carry no clicks, no click-through rate, no position and no query dimension. The data was largely already present in the overall report; what shipped was a view, not a new measurement. It is a second witness by design, and a deliberately narrow one. This is also the honest context for arguments about what an agentic visit is actually worth when it never produces a click: the platform has chosen which half of that question it will help you answer.

What it will not restore is ordinality. Nothing in the requirement produces a ladder, and no regulator is going to mandate one. Plan for a world with better attribution and no ranking, rather than for the return of a scoreboard.

8. Worked example: Bartholomew Dane, Nottingham

A hypothetical but deliberately ordinary case. Bartholomew Dane is a Nottingham distributor of laboratory and analytical equipment, £12.4M turnover, selling to universities, contract research organisations and QC labs. It had run a £6,500-a-month agency retainer since 2023, with a performance clause paying an extra £1,500 in any month where at least 25 of 60 tracked terms held a top-three position. Both sides read the number off the agency’s rank tracker.

January 2026 — the number stops meaning anything

Thirty-one of the 60 terms held top three, so the bonus paid every month of the quarter. Organic revenue over the same period fell 9% year on year. The finance director asked the obvious question and nobody could answer it, because both parties were reading a measurement whose performance content had quietly drained out of it while its arithmetic kept working.

March 2026 — the replacement fails the same test

They trialled two AI visibility platforms and ran both against the same 40 prompts for a month. One reported 34% citation share; the other reported 11%. Same prompts, same period, same brand. Neither could be reconciled with the other, because neither published its panel composition, run frequency or prompt-selection method. The agency preferred the first platform. The client preferred the second. Nothing in either product could adjudicate that, and both were honest — they were sampling different systems in different ways.

April 2026 — the sort

They put all 14 reporting metrics through the three witnesses. Three came out as first witnesses: named placements at live URLs, two trade-register listings, and one published dataset the company itself maintained. Six were second witnesses — Search Console, analytics, CRM — all reported with no written method. Five were third witnesses, including both citation-share figures and the blended visibility score that had been the headline slide since January.

May–August 2026 — the rewrite

The rank clause was removed. Performance moved onto three things. Named placements against an agreed publication tier list, each verified by loading the URL and archiving it. A quarterly citation sample run by the client to a written protocol — 60 prompts, five runs each, three engines, results shared in full including the misses. And qualified enquiry volume from a shared CRM view with agreed stage definitions. The first is a first witness; the other two are second witnesses promoted as far as a written method can promote them.

By August: 26 placements delivered against a target of 24, referring domains up 38, the quarterly citation sample moving 19% to 31% across three runs, qualified enquiries up 17%, and organic revenue back to flat year on year.

The internal effect was larger than the external one. Once the bonus stopped depending on a number the agency produced, the monthly call changed character — the first twenty minutes had previously been spent litigating tracker configuration, and that argument simply had nowhere to go. Both sides described the new arrangement as slower and less satisfying to report. It was also the first time in three years that the two organisations were looking at the same evidence.

What went badly

The quarterly sample cost roughly nine hours of a client analyst’s time per run and was nearly dropped in round two. The tier list took six weeks of argument because “editorially independent” had to be defined, and two of the agreed publications later ran a near-identical piece for a competitor. Removing the rank clause cost the agency its bonus in two months where rankings genuinely improved, which they resented and which took a difficult conversation to settle.

And the first-witness category proved less permanent than the framework implies. Of the 26 archived placements, three pages were substantially rewritten within four months and one publication dropped the link entirely during a CMS migration. Public verifiability turned out to need a quarterly re-check, which nobody had budgeted for — and which is now the least glamorous line in the retainer.

9. Where this argument could be wrong

The hardest counter is that this is nostalgia wearing an analytical costume. Rank was never a good settlement metric: it was gameable, localised, personalised, and the industry spent a decade saying so out loud. Citation-share estimates are arguably more honest, because they report uncertainty instead of concealing it behind a simulated location. On this reading, what has happened is that a bad number has been replaced by an honest one, and the discomfort is just adjustment.

Most of that lands, and the honest response is to concede it rather than argue. A shared flawed number can settle a dispute; a private accurate one cannot. That is the whole claim, and it does not require rank to have been any good. The second half of the counter — that contracting on links recreates the incentives that produced the spam era — is the part that genuinely bites, and it is bounded only by writing clauses about specified placements rather than counts. A publication that has spent years arguing about disavowal and manual actions is in no position to pretend link volume is a clean KPI, and this one is not pretending.

Two further bounds. Business outcomes remain the correct ultimate settlement metric wherever attribution genuinely supports them; the three witnesses is a framework for the intermediate measures that fill the eighteen-month gap between work and revenue, not a replacement for revenue. And the argument only applies where two parties need to agree — a solo in-house team with no external contract has a measurement problem, not a settlement one, and should read the third-witness numbers gratefully.

A fourth bound is worth stating because it cuts against the tidy version of this argument. Verifiability and importance are not the same property, and elevating what can be checked risks quietly demoting what matters. Brand demand, category understanding, the slow accumulation of the reputation a model absorbs during training — none of these produce a document anybody can load, and all of them plausibly outrank a placement count in real commercial consequence. The three witnesses tells you what a number is allowed to settle. It does not tell you what to care about, and treating it as though it did would reproduce the original error of mistaking the measurable for the important. That error is what a decade of PageRank sculpting was made of.

Three developments would materially damage the case. An engine publishing per-query citation logs to verified site owners, which would make citation a first-witness fact overnight. An independent, non-commercial body producing an audited citation index with a fully reproducible method. Or evidence that the vendor panels are converging — the current data points the other way, but convergence is exactly what you would expect if the underlying systems standardised, and it should be checked annually rather than assumed away.

10. The Monday checklist

  • List every metric in your current reporting pack and mark each one first, second or third witness. Do it before you look at the numbers themselves.
  • Find any third-witness number that appears in a contract, a bonus calculation or a board commitment. That is your immediate exposure.
  • For every second-witness figure you intend to keep, write the method down: definitions, date ranges, filters, who runs it. Circulate it and get it agreed before the next reporting cycle.
  • Ask each of your visibility vendors, in writing, for panel composition, run frequency and prompt-selection method. File the answers — and file the silences.
  • Build a placement tier list with your counterparty and define “editorially independent” in words you would both accept in a dispute.
  • Archive every placement URL on the day it goes live, and put a quarterly re-check in the calendar. Assume roughly one in six will move, change or vanish within a year.
  • Design your own citation sample rather than buying one: fixed prompt set, repeated runs, engines named, misses reported. A crude method you control beats a sophisticated one you cannot inspect — and it is cheap next to the platforms selling the same sampling.
  • Re-plan the quarter around acquiring verifiable placements rather than defending positions, using whichever acquisition tactics fit your market, and benchmark against the current published link and citation data rather than a vendor score.

The uncomfortable truth underneath all of this is that the search industry was never as measurable as it believed; it was adjudicable, which is a different and rarer thing, and it was adjudicable because a third party with no interest in the outcome published a number for free. That arrangement has ended. What replaces it is not a better metric — it is the older, slower discipline of producing evidence somebody else can check. Which is, awkwardly for anyone hoping this was a reporting problem, the same discipline as earning the placements themselves.

Leave a Reply

Your email address will not be published. Required fields are marked *

Query Fan-Out Previous post Query Fan-Out: How AI Mode Splits One Question Into Many (and Cites Accordingly)
Zero-Click Traffic Model Next post Rebuilding a Traffic Model for the Zero-Click Majority