AI Mode Source-Diversity

AI Mode’s Source-Diversity Bias: Why It Cites More Domains — and How to Be One

TL;DR

— AI Mode does cite far more domains than AI Overviews. But that breadth is spread across queries, not stacked inside answers. Per answer, you are still competing for a handful of slots.

— Those slots are not awarded to the best pages. Google Research describes this class of retrieval as set-valued and non-decomposable: the set is scored as a set, with an explicit diversity term whose job is to stop the system returning near-duplicates.

— So your real competitor is not the page ranked above you. It is whichever source already said what you were about to say. Relevance gets you considered. Non-redundancy gets you selected.

— Two instruments below: the Role Inventory, which shows which evidential roles an answer fills and which sit empty, and the Substitution Test, which converts your redundancy into odds.

— UK-specific: the CMA’s publisher conduct requirement of 3 June 2026 lets UK publishers withdraw from AI Mode at page level. Every withdrawal empties a role. Britain is the only market where the source set is being thinned by regulation, and the vacancies are dated.

The claim is true. The conclusion drawn from it is not

Start by conceding everything, because the measurement work here is good. Otterly’s comparison of the two Google surfaces found AI Mode drawing on 3,621 unique domains against 615 for AI Overviews, producing 30,353 citations against 2,498 — and triggering on 100 of 100 test queries where AI Overviews appeared on 49. Ahrefs, working across 540,000 query pairs, put AI Mode at roughly nine unique domains per query against 7.7 for AI Overviews. Moz, across nearly 40,000 queries, found only 12% of AI Mode citations matching a URL in Google’s organic top ten. A December 2025 arXiv analysis found 37% of domains cited by AI search engines were absent from conventional results entirely.

Every one of those numbers points the same way, and the field has drawn the same conclusion from them: the door is wider, so walk through it. Publish for the long tail. Cover more sub-queries. Being one of many is easier than being first of ten.

That inference does not survive contact with how the set is actually built. The breadth is real, but it is breadth across queries, not within them. Attrifast’s 1,200-prompt benchmark put the median number of unique domains in a single answer at 2.4 for Gemini, 3.1 for ChatGPT, 3.6 for Claude and 6.4 for Perplexity. Thousands of domains, three or four at a time. What the diversity statistics measure is how much the identity of those three or four changes as you move from question to question — which is variance, not generosity.

What the aggregate numbers are actually counting

The studies disagree wildly on concentration, and the disagreement is informative. Goodie’s analysis of 58.6 million citations put Wikipedia — the single most-cited domain — at 3.4% share, which reads as a flat, open long tail. Everything-PR’s synthesis of six studies covering more than 680 million citations put the top fifteen domains at roughly 68% of all citations, which reads as an oligopoly. Both can be right, because they are counting different things: one is measuring the corpus, the other is measuring the slots. A corpus-level statistic tells you how many domains the system has ever used. It tells you nothing about your odds in the answer you care about.

This matters for anyone allocating budget on the strength of these figures. If you want the numbers themselves, our running index of link building and citation statistics tracks the surviving ones. But the number to act on is not the count of domains cited. It is the count of domains cited in the answers your buyers see, and that number has barely moved.

What is source diversity in AI Mode? It is a deliberate property of the selection step, not a side effect of having a big index. The system is optimising for a set of sources that jointly covers a question without repeating itself, which means each additional citation is judged on what it adds to the ones already chosen — not on its own merit in isolation.

A set is not a list, and that changes who you are competing with

In February 2026 a team at Google Research and the University of Illinois published a paper on fan-out retrieval — the technique that decomposes one question into many sub-queries — which describes the problem class with unusual precision. Retrieval systems, they write, are increasingly expected to return sets of results rather than a single best match, and set-valued objectives are typically non-decomposable. Quality is defined by set-level properties: diversity, intent coverage, complementarity, coherence.

Non-decomposable is the load-bearing word. It means the quality of the set cannot be recovered by adding up the quality of its members. There is no score your page carries into the room. Your contribution is defined only against what else is in the set — which changes per query, per sub-query, and per run.

The paper’s reward function makes the mechanism explicit. It combines three terms — groundedness, alignment with the original question, and diversity — weighted 0.6, 0.2 and 0.2 in the authors’ default configuration, with diversity measured by the Vendi Score, a published metric for how much semantic spread a set contains. Then comes the finding that should reorganise how link builders think about this surface. Without the diversity term, the reward is trivially gamed: trained on groundedness and alignment alone, the model collapses faster, because the policy simply repeats paraphrases of the original query. Their conclusion is that diversity and alignment act as mutual counter-anchors.

Read that again with your content strategy in mind. The diversity term is not a quota for newcomers. It is a guard against paraphrase. It exists because a system optimising purely for relevance produces near-identical results, and near-identical results are worthless in an answer. Which means the thing that gets a source excluded is not being worse. It is being the same.

Why the second slot does not go to the second-best page

Under a set-level objective with a redundancy penalty, selection is marginal and sequential. The first slot goes to the strongest available source. The second does not go to the runner-up; it goes to the source that adds most given the first is already there. If the runner-up is a well-written page covering the same ground as the winner, its marginal contribution is close to zero and it is passed over for something worse but different.

An honest boundary: that paper benchmarks fashion and music corpora, not Google Search. It is not a description of AI Mode. What it is, is Google’s own researchers stating the objective structure of the retrieval class that AI Mode’s fan-out belongs to — and Google’s separate patent for the fan-out step is titled, in its own words, prompt-based query generation for diverse retrieval. Nobody outside Mountain View can confirm the weights. The structure is not in dispute.

Key takeaway

Stop asking whether your page is good enough to be cited. Under a non-decomposable objective, that question has no answer. Ask instead: given the sources already in this answer, what does mine add? If the honest reply is a second opinion on the same point, you are competing for a slot the system is specifically designed not to award twice.

Instrument one: the Role Inventory

Diversity is not enforced across domains. Domains are the unit you can count, which is why every tool reports them, but a coverage objective partitions on something else: the evidential role a source plays in supporting the answer. A register and a review site are not two domains competing for two slots. They are two different kinds of evidence, and an answer needs both.

The Role Inventory below is the first of this article’s two instruments. Take twenty real buyer prompts, run each in AI Mode with your market’s locale set, capture every citation, and classify each one by role rather than by domain authority. The column that decides your strategy is the last one.

Evidential roleThe question only this role answersWho fills it in a UK answer setCan your own domain fill it?
Canonical definitionWhat is this thing, stated neutrally?Wikipedia, BSI, government guidance, standards bodiesNo. Permanently occupied.
Primary specificationWhat exactly does this product do, at what spec, at what price?The vendor. Manufacturer sheets, docs, product pagesYes. The one role you own outright.
Independent comparisonWhich of these is better, and for whom?Comparison sites, trade titles, Which?, review platformsNo. Self-comparison is read as testimony and discounted.
Verified statusIs this firm real, registered, accredited, insured?Registers and regulators: Companies House, MCS, Gas Safe, FCA, TrustMarkNo, by construction. You are listed in it; you cannot publish it.
Measured evidenceWhat happened when somebody actually measured it?Test labs, universities, first-party studies carrying a method and a dateYes, but only if you genuinely ran the measurement.
Practitioner testimonyWhat was it like in practice, from someone with nothing to sell?Forums, Reddit, trade communities, professional bodiesNo. Cannot be manufactured without serious risk.
Dated event recordWhat changed, and when?National and regional press, trade news, regulator announcementsNo. Earned only, and it decays with the news cycle.

Colour band shows whether the role is addressable from your own estate: green = yours to take, amber = earnable but perishable, red = structurally closed to you.

How to read your own inventory

Three states matter. A filled role has one stable incumbent and is not worth attacking. A contested role has three or more sources doing the same job, which is a substitution war you are probably already losing without knowing it. An empty role is one where the answer visibly needed something the engine could not find — a hedge, a vague generalisation, a claim with no citation attached. Empty roles are where unremarkable pages get cited above excellent ones.

For example, a UK commercial insurance broker running this exercise found six roles across their prompt set, five occupied and one — the dated event record for a regulatory change eight weeks old — held by nobody, because the trade press had covered the announcement but not its practical effect. That gap was worth more than another comparison page. Related: our note on how engines weigh product recommendation factors covers why some roles carry disproportionate weight in commercial answers.

Instrument two: the Substitution Test

The Role Inventory tells you where the gaps are. The Substitution Test tells you what your current position is worth, and it produces a number you can put in a plan.

THE SUBSTITUTION TEST

1. The deletion question. Remove your citation from the answer. Does any sentence lose its support? If the answer holds together untouched, you were not evidence. You were decoration.

2. The class count. Count every other cited source in that answer that could have answered the same sub-question. Call the total, including you, n. Before quality tie-breaks, your odds of holding that slot on any given run are roughly 1 in n.

3. The doubling question. Would the answer still need you if the incumbent’s page were twice as good? Yes means you are a complement, and you are priced on scarcity. No means you are a substitute, and you are priced on luck.

The arithmetic that follows. Moving from a class of eight to a class of two takes your per-run odds from about 12% to about 50%. Writing a better page and moving from eight to seven takes them from 12% to 14%. You cannot out-quality your way out of a substitution class. You can only leave it.

The evidence that most cited sources are substitutes

Ahrefs’ comparison of the two Google surfaces is, on this reading, the most important dataset published in 2026 and it was reported as a curiosity. Across 540,000 query pairs, AI Mode and AI Overviews cited the same URLs only 13.7% of the time — yet their answers were semantically similar in roughly 86% of cases. The same conclusions, from almost entirely different pages.

That is the definition of a substitute. If swapping out 86% of the sources leaves the answer unchanged, those sources were interchangeable, and interchangeable sources are selected by something very close to a coin toss. SE Ranking’s AI Mode study found only 9.2% of URLs matched across three runs of the same query on the same day. SparkToro and Gumshoe, across 2,961 prompts with 600 volunteers, put the chance of two responses returning the same brand list at under one in a hundred.

Why does my citation share change every time I measure it? Because you are inside a substitution class. Sources that uniquely fill a role are cited every run; sources that duplicate another source’s role are sampled. Volatility is therefore not a defect in your monitoring tool — it is a measurement of your redundancy, and it is the cheapest diagnostic on this list. Stability, not share, is the metric that tells you whether you own something.

Why competent SEO now maximises your chance of being deduplicated

The standard playbook is convergent by design. Study the pages that rank, cover what they cover, close the gaps, be comprehensive, match the intent. Executed well, it produces a page maximally similar to the incumbent — which, under a redundancy penalty, is the single most reliable way to be left out. Comprehensiveness is a redundancy strategy wearing a quality costume.

The fan-out advice compounds it. Told that one query becomes eight, teams build one pillar page addressing all eight sub-queries. But three or four of those sub-queries — comparison, verification, constraint — cannot be answered from your own domain at all, because an answer to is this firm accredited or which of these is better carries no weight when the firm itself supplies it. A page that answers all eight is not eight opportunities. It is one opportunity plus seven claims the engine will discount as self-testimony.

For example, a Leeds recruitment software vendor audited a 6,000-word pillar page built exactly this way and found it cited for none of its eight sub-queries. The two commercial ones were held by a professional body and a comparison site — a pattern that recurs across recruitment and HR technology, where the verification role is owned outright by institutes nobody can displace.

This is also why chasing featured-snippet-style coverage transfers badly. Snippet optimisation rewarded being the cleanest statement of the consensus. Set selection penalises exactly that, because the consensus is already in the set.

Key takeaway

Before commissioning another comprehensive guide, run the deletion question against your last three. If removing them would leave every answer in your category intact, the correct next investment is not a fourth guide. It is one document that says something no other cited source in your market is positioned to say, and one placement in a role you cannot occupy yourself.

The UK case: the only market where roles are being emptied by law

Everything above applies everywhere. What follows applies only in Britain, and it has a date on it.

On 10 October 2025 the Competition and Markets Authority designated Google as holding strategic market status — SMS, a formal finding of entrenched market power — in general search and search advertising, on the basis that Google handles more than 90% of UK searches. On 28 January 2026 the CMA consulted on four conduct requirements, closing on 25 February with more than sixty published responses. Two decisions followed.

3 June 2026 — the publisher conduct requirement. Google must give publishers a control determining whether their content is used in AI-powered search features, explicitly including AI Mode as well as AI Overviews. The CMA adopted publisher recommendations on fine-tuning and, critically, page-level controls. The regulator described it as a world first.

17 June 2026 — fair ranking and data portability. Google must rank organic results using objective and non-discriminatory criteria, including in search generative AI features, give greater transparency to businesses, and provide notice before ranking changes. It has six months to implement fair ranking and three for data portability.

Why a page-level opt-out is a link building event

Read the publisher control through a set-level objective and it stops being a publisher story. When a page is withdrawn, the system does not lose a page. It loses whatever role that page was filling — and a coverage objective must still fill the role. The slot does not close. It reopens.

Now ask who is most likely to use the control. Ad-funded national news, whose economics are worst served by uncompensated summarisation, and who happen to occupy two of the seven roles in the inventory above: the dated event record and, through their reviews and comparison desks, independent comparison. Meanwhile the roles held by UK registers, regulators, trade bodies and professional institutes — verified status, canonical definition — are held by organisations with no advertising model, no compensation grievance and no reason to withdraw. Those stay locked. The vacancies open in exactly the roles that are earnable.

One distinction is worth holding onto here, because it gets collapsed constantly. The CMA’s control governs whether content is used in AI search features — a retrieval-side permission, taking effect the moment it is set. It is not the same decision as whether a crawler may use the material as training data, which is governed separately and, across Europe, sits alongside the transparency duties in the EU AI Act. A publisher can withdraw from AI Mode and stay in the training corpus. When auditing which of your own placements survive into next year’s answers, treat the two permissions as separate line items on different clocks.

There is a second UK-specific asymmetry underneath it. Profound’s data puts .co.uk at roughly 2.16% of ChatGPT citations globally, while Oxford Internet Institute work has found these systems systematically favouring English-speaking regions. A thin national supply inside a favoured language is the definition of a market where role vacancies are easy to fill — the opposite of the position UK firms face on generic global queries. If you operate across borders, our guide to running link acquisition in more than one market covers how badly this asymmetry travels.

The regulatory mismatch worth understanding

The fair ranking requirement is written in the vocabulary of ranking: objective, non-discriminatory criteria for ordering results. A non-decomposable set objective has no ordering of that kind to be fair about. The same page can be included in one answer and excluded from the next with no change in its own properties, because the exclusion was caused by something else being in the set. Transparency about criteria therefore does not make membership predictable — and the thirty-day notice that will genuinely help with algorithm updates has no purchase on run-to-run variance at all. This is an observation about mechanism, not a judgement on the policy, and the compliance deadline in December 2026 is the first real test of it.

What you actually buy when you buy non-redundancy

Here is the awkward implication for content teams. Non-redundancy is not a writing problem. You cannot become a different kind of source by writing more carefully. A trade register is non-substitutable because of what it is, not because of how it words things. Which means the roles you cannot fill from your own domain have to be occupied by someone else on your behalf — and that is the definition of link building, whatever the field is currently calling it. If you need the ground floor of that argument, our primer on what link building is for still holds, but the selection criteria have moved.

Muck Rack’s analysis of over a million cited links found 89% of ChatGPT citations originating in earned media rather than owned, paid or social content. That figure is usually quoted as an argument for digital PR volume. It is better read as an argument about roles: earned media dominates because four of the seven roles admit only third-party evidence by construction.

Commission by role, not by metric

The practical change is to the brief. Stop specifying referring-domain counts and authority thresholds; start specifying which role a placement is meant to fill and which sub-question it must be able to answer. A county-level scheme listing with a Domain Rating of 24 that uniquely answers is this provider approved here outperforms a DR 78 guest post that repeats the category explainer, and it does so on every run rather than one in eight.

That reprices most of the standard tactics. Guest posting and paid link insertions into existing posts almost always land in the explainer role, the most contested one on the board. Sponsorship placements and local citation entries are usually dismissed as low-value, but frequently sit in verified status — a role nobody can take from you. Journalist-request platforms and news-hook outreach reach the dated event record, which is the role the UK opt-outs are about to vacate. Launch-platform placements reach practitioner testimony. The tactic list has not changed; the reason for choosing among them has.

One warning on the measurement side. Because role occupancy is stable and share is not, any tracker sampling a prompt once will misread both. Sample the same prompt repeatedly and report how often you hold a role, not what percentage of citations you took. Our comparison of monitoring and link building tools flags which platforms expose per-run data rather than an averaged score, and our note on measuring entity authority covers the adjacent problem of attributing this to a brand rather than a URL.

Where this argument is weakest

The strongest objection is not that the mechanism is wrong. It is that the mechanism is being turned down, and that this whole framing describes a regime that is closing.

The evidence for it is good. Evertune ran 99,000 prompts across Gemini and Google AI Mode over April and May 2026, on 25 days in each month. Between late April and the end of May, the number of unique URLs cited fell 33% on Gemini and 59% on AI Mode. Their read: the sources that still get cited now carry more weight than they did a month earlier. The Google Research paper supplies a candidate reason, noting that stronger reward optimisation typically reduces output entropy and encourages mode collapse. If the diversity weight is falling, the argument runs, AI Mode is converging on a concentrated authority ranking and the old playbook — be the biggest, best-linked source in your category — simply wins again.

Taken seriously, that concedes a great deal. It should be bounded four ways.

First, it sharpens the thesis rather than refuting it. If the set shrinks from nine domains to four, substitution classes do not shrink with it. Your class size stays the same while the slots halve, so your odds inside a contested role get worse, not better. Non-redundancy becomes more decisive as the door narrows, not less.

Second, the concentration benefited nobody. Of more than 20,000 domains AI Mode cited in both months, exactly one — YouTube — gained even a single percentage point of source share. A tightening that produces no winners is a threshold move, not a return to authority ranking, and a threshold you have not been told the height of is not something you can optimise toward.

Third, it is one tracker over five weeks. Otterly, Ahrefs, Victorious, Moz and SE Ranking measured over different windows with different methods and disagree on levels by wide margins. Directional readings from single windows have been wrong repeatedly on this surface; averaging across them is worse than useless.

Fourth, the recommended actions are invariant. Occupying an under-filled evidential role is correct whether the answer holds nine sources or three. Nothing in this article asks you to bet on the diversity weight staying where it is.

Three things would falsify the argument, and any reader can run the second inside a month. One: panel data showing citation stability is uncorrelated with how many co-cited sources answer the same sub-question. Two: a controlled test in which a page written to duplicate an incumbent’s coverage is cited as often as a matched page written to fill an empty role. Three: Google documenting that AI Mode scores each source independently of the set. Until then, the burden sits with the coverage-everything playbook, which has never explained why 86% of answers survive an 86% change of sources.

Worked example: Wexbury Thermal, Banbury

Wexbury Thermal is an invented but specified case: a heat pump installer in Banbury, Oxfordshire, £6.8M turnover, 44 staff, working across Oxfordshire, Buckinghamshire and Northamptonshire. Domestic retrofit, mostly grant-assisted. Their category is unusually rich in UK institutional sources, which makes the role structure legible.

February 2026 baseline. Sixty buyer prompts run in AI Mode with UK locale, three runs each — 180 runs. Wexbury appeared in 21% of runs, which their agency reported as reasonable progress. Classified by role, the picture changed. Every appearance sat in one role: the category explainer. Ten other sources answered the same sub-question in the same answers, among them the Energy Saving Trust, Which?, two national installers and three trade magazines. Substitution class of eleven; expected odds roughly 9%. Their observed 21% was flattered by a handful of runs where a second page was pulled in.

What the inventory exposed. Three roles were thin or empty across the prompt set. No cited source gave installed costs by property archetype for their region — answers hedged with national averages. No source verified scheme eligibility at county level. And nothing at all occupied measured post-install performance: what SCOP, the seasonal coefficient of performance that determines running cost, systems actually achieved against what was quoted.

Sixteen weeks of work. They published a dated regional cost study drawn from 340 of their own completed installations, broken down by archetype, with method stated — first-party operational data they already held, not a commissioned survey. They published measured SCOP against quoted SCOP for 118 systems across two heating seasons. They took listings in two trade registers and a county retrofit scheme. And they commissioned or earned fourteen placements chosen by role rather than authority score: a county council retrofit resource page, two regional papers covering a specific local scheme, a further education college’s training page, a consumer trusted-trader profile, and three trade titles running features on measured performance.

At twenty weeks. Their share in the explainer role was unchanged at roughly 9% — exactly as predicted, because they never left that class. But they appeared in the cost-evidence role in 61% of runs and the performance role in 38%. Blended run-presence went from 21% to 47%. Enquiries attributable to AI surfaces rose from 9 to 24 a month. The number that mattered most to them was neither: across 36 repeat runs, the cost study was cited 34 times and the old explainer page 4. Same company, same domain, same authority profile — a 94% hit rate against 11%, decided entirely by which role each document occupied.

What went wrong. Publishing installed costs by archetype handed competitors a price sheet: two repriced within six weeks and a national undercut them on precisely the archetypes they had published. Three of the fourteen placements produced no citation in twenty weeks because the role was already held by a fifteen-year-old page nobody was going to displace. The cost study needed re-dating quarterly at about fourteen hours a time or citations decayed. And in month five the county council restructured its resource page and dropped their mention without notice — the standing hazard of an asset that lives on someone else’s server, discussed further in our note on defending an earned link profile.

The Monday checklist

Eight things, in order. The first four take a week and cost nothing but analyst time.

1. Pick twenty real buyer prompts. Run each three times in AI Mode with your market’s locale set. Capture every citation, with the run number.

2. Classify every citation by evidential role, not by domain or authority score. Use the seven roles above until your category demands an eighth.

3. Mark each role filled, contested or empty. Contested means three or more sources doing the same job.

4. Run the deletion question against your own citations. If the answer survives without you, record it as decoration, not evidence.

5. Count your substitution class in every role you appear in. That number, not your citation share, is your position.

6. Identify the one measurement only you could publish — operational data you already hold, with a method and a date. Ship it before the next guide.

7. Rewrite one outreach brief to specify a role and a sub-question instead of a referring-domain target, and place against it.

8. Re-run the twenty prompts at week twelve and report stability — how often you hold each role — never averaged share.

The uncomfortable part of all this is that it demotes quality from a cause to a qualifier. Being good gets you into consideration; it does not get you into the set. What gets you into the set is being the only source positioned to answer a question the answer cannot avoid asking — and for most businesses, most of those questions will always be answered by somebody else’s page. Which is the same conclusion the field reached about links twenty years ago, arrived at from an entirely different direction. If you want the tactical map for acting on it, our index of link building strategies is organised by what each approach can and cannot reach.

Leave a Reply

Your email address will not be published. Required fields are marked *

Multi-Turn Query Chains Previous post Optimising for the Follow-Up: Multi-Turn Query Chains and Sustained Citation
SERP-Less Audit Next post The SERP-Less Audit: Measuring Presence When There is No Results Page