TL;DR
Model collapse is a real result about a training regime almost nobody operates. It requires each generation to replace its predecessor’s data. Accumulate instead and the error stays bounded.
Collapse deletes the tails of a distribution. It does not read bylines. A human-written page in the fat middle is the material collapse cannot damage and retrieval has no reason to fetch.
What decides whether an engine looks at all is its own uncertainty. Unknown brands force a search; familiar ones are answered from memory. Live search fires on roughly a third of ChatGPT prompts, down from about 46% in late 2024.
So the operative property is not human authorship. It is unpredictability: whether the rest of the corpus could have written your sentence without you.
And novelty is consumed by being published. A citation is a receipt for the engine’s ignorance, and the receipt expires.
1. What model collapse actually says
What is model collapse?
Model collapse is the degradation that occurs when a generative model is trained on data produced by earlier generative models, generation after generation. It is a statement about a recursive sampling process rather than about content quality: the model’s picture of the world drifts, and the rare events at the edges of the distribution disappear first.
The formal result is Shumailov and colleagues in Nature, July 2024. Train a model, sample from it, train the next model on those samples, repeat. Two things happen in sequence. In early collapse, estimation errors accumulate and the model drifts from the true distribution while still producing fluent output. In late collapse, variance shrinks towards zero and the model converges on something close to a single point: in the Gaussian illustration the estimated covariance goes to zero while the mean wanders off.
The tails go first, and the reason is arithmetic
Each generation draws a finite sample from its parent, and low-probability events are precisely what a finite sample is likely to miss. Miss a rare event once and it is absent from the next model’s training data; absent once, it is gone. Alemohammad and colleagues found the same shrinkage in image models, naming the failure mode after autophagy.
Collapse is indifferent to who wrote a document. It deletes improbable material. A human-written page saying what a thousand other pages already say sits in the fat middle and survives every generation untouched — and is worth nothing to a retrieval system for the same reason. Being safe from collapse and being invisible to retrieval are one property seen from two ends.
The condition everybody drops
The Nature result assumes each generation replaces its predecessor’s data. Gerstgrasser and colleagues asked the obvious follow-up: what if you accumulate — keep the original real data and add each generation’s synthetic output alongside it? Under accumulation the test error stays bounded rather than growing without limit. Alemohammad’s taxonomy makes the same cut with three loop types: fully synthetic loops collapse, augmented and accumulating loops with enough real data do not. A separate line of work arrives via verification: filter with a good-enough verifier and recursive training stabilises. None of that is exotic — it describes what a competent lab already does, and Microsoft’s Phi work showed curated synthetic data taking a 1.3-billion-parameter model to parity with far larger ones.
Key takeaway. Collapse is conditional on replacement. The useful question is not whether models degrade but which material they lose — and the answer is the improbable material, regardless of who produced it.
2. The forecast the industry actually bought
The commercial version runs like this. The open web is filling with machine-written text. Models trained on it will degrade. Human-written material will become scarce, scarce inputs get repriced, so keep publishing human content and the market will come to you. Every step has a number attached and the numbers are real. Roughly half of newly published articles are primarily machine-written, a share flat for about five quarters. Ahrefs, across 600,000 pages, found 86.5% contain some machine-generated text but only 4.6% are wholly machine-written — which quietly kills the binary the debate is conducted in. Originality.ai put AI pages in the top twenty at 19.56% at their July 2025 peak, and NewsGuard’s count of AI-generated news sites went from 49 to 1,271 in two years. Epoch AI’s estimate of the effective stock of quality-adjusted public human text, around 300 trillion tokens, implies full utilisation between 2026 and 2032, with the front of that window since revised out toward 2028. The supply-side picture for AI training corpora has not changed much since; the interpretation has.
The premises are sound. The conclusion does not follow, for two reasons.
Even where it bites, collapse deletes tails, not humans
No step in the mechanism inspects provenance, because there is nothing in a document to inspect. If a claim is well represented across a corpus it survives; if it is rare it goes, and it goes whether a person or a script produced it. A strategy that reliably produces conventional prose in a human voice is manufacturing the one class of material collapse cannot touch and retrieval has no reason to fetch.
The repricing happens at a counter you cannot reach
Where scarcity has actually been priced, it has been priced by the corpus and not by the page. The visible transactions are licensing deals over whole archives and whole platforms. Those cheques are written against volume, rights and warranties, they arrive on a model-release cycle, and they carry no line item for anybody’s blog post. A publisher waiting for the collapse thesis to reprice their content is waiting on a market in which they hold no seat and can obtain none.
Key takeaway. The prediction confuses two scarcities. Human text may well become scarce as a bulk input. Your particular page becomes valuable only if it says something the corpus cannot say without it. Only the second is under your control.
3. The variable that actually operates is predictability
Does it matter to an answer engine whether a human wrote the page?
Not directly. No production retrieval system reads a provenance signal, and the studies most often cited for a human advantage measured writing style, not authorship. What a retrieval system responds to is whether a document adds information the rest of its corpus does not already contain.
That property has a precise name in information theory, and borrowing it turns a vague argument about quality into something testable. The information carried by a statement is inversely related to how predictable it was: a sentence the corpus could have generated without you carries close to nothing. This is not a metaphor. It is a description of the objective these systems are trained on — a model is a machine for predicting the next token, so publishing text it can already predict is, mechanically, publishing nothing.
Human authorship is neither necessary nor sufficient. A person writing the modal paragraph on a well-covered subject adds zero; a twenty-line script joining two public registers nobody has joined adds a lot. The observed correlation — human-written pages rank and get cited more often — runs through the production process, not through a detectable signal in the document. The acts that generate new claims (measuring something, asking somebody, going somewhere) still require people. Human involvement is a symptom of novelty in 2026, not its cause, and treating a symptom as a cause is how a discipline ends up buying bylines instead of vantage points. The case for why experience alone does not command a premium turns on the same distinction.
We industrialised the production of predictable text
Almost every core technique in the discipline is a device for finding the centre of a distribution. Keyword research selects the phrasings most people already use. Briefs assembled from the top ten regress content towards the mean of what already ranks. Comprehensiveness — cover every subtopic your competitors cover — is modal coverage by definition. Even the well-meaning parts of standard content and link strategy pull the same way. A team writing the twelfth guide to a compliance regime, built from the eleven ranking above it, has produced a document whose defining property is that it was predictable from the eleven.
That was survivable when the reward was a ranking and the rivals were other pages. It stops being survivable when the reward is a fetch and the rival is the model’s own weights. The same trap sits under most featured-snippet-shaped content: you win by saying the expected thing in the expected form, which is precisely what a model no longer needs a source for.
4. The Four Sources of Net-New
“Net-new” does all the work in “net-new human knowledge”, and it is not one thing. Four structurally distinct reasons make a claim unpredictable, and they differ in cost, in lifespan, and in whether anybody else can carry them. Treating them as one budget line is the most expensive planning error in this area.
| Source of novelty | Why the corpus cannot predict it | What ends it | Who else can carry it |
| Time (recency) | It had not happened at the last crawl. Novel because of the clock, not the work. | The next crawl. Days to weeks. | Everyone, immediately. Recency is non-excludable. |
| Standing (access) | You are the only party observing the population. It cannot be inferred from public material. | Losing the access, or a rival building a comparable vantage. Years. | Only a party holding their own vantage, which is why joint studies exist. |
| Labour (computation) | The inputs are public but nobody has performed the operation. | Anyone willing to repeat the work. Weeks to months. | Anyone with an analyst and the same public inputs. |
| Authorship (constitutive) | The fact exists because you stated it: your prices, terms, policies, method. | You changing it. | Nobody. There is nothing to corroborate. |
Green marks the row producing durable, transferable value. Amber produces real novelty on a short lease. Red produces novelty everybody gets at once, or that nobody else has reason to repeat.
Time is the row the field over-buys. Reaction pieces and news-pegged campaigns are unpredictable at publication and worthless a fortnight later, because recency is non-excludable — everybody receives the event at once. This is not an argument against speed, but against booking speed as an asset.
Standing is the strongest and most expensive. If you run sixty refrigerated vehicles, you observe a population no competitor, consultancy or model can reach without building the same fleet. Standing is not defeated by effort — a rival cannot out-work you into your own meter readings. It is also the row most often thrown away by accident, because the data constituting it sits unretained as an operational by-product.
Labour is the underpriced one. A great deal of unpredictable material comes from entirely public inputs, simply because nobody has performed the operation: a join across two registers, a series nobody assembled, a re-derivation of a widely quoted figure that turns out to trace to a single 2019 press release. Its weakness is symmetrical to its strength — what one analyst does in a fortnight, another repeats in a fortnight.
Authorship is the trap. Facts about yourself carry maximum surprisal and minimum weight. Nothing predicts your pricing page, so an engine needing your prices must fetch it. But a constitutive fact says nothing about the world beyond your own decision, making you an authority on nothing except yourself. This is why the exhaustive “our approach” estate fails against commercial prompts despite being, strictly, entirely original — and why entity authority has to be measured off-domain or not at all.
5. Retrieval is triggered by the engine’s ignorance, not your quality
When does an answer engine actually go and look?
Less often than most content strategies assume, and the decision is made before your page is anywhere near the process. A live search fires on roughly a third of ChatGPT prompts. Everything else is answered from the weights.
The measurements cluster tightly. Nectiv put search frequency at 31% of prompts; a 391-query test through the web interface found 42%; a 2026 compilation put it at 34.5%, down from around 46% in late 2024. The direction and order of magnitude are consistent, and they sit alongside the wider AI citation and referral figures.
The trigger rate is not distributed randomly across query types, and the pattern is this article’s argument restated as telemetry:
- Definitional queries about well-covered subjects almost never trigger a search. The model has the answer and knows it.
- Comparison queries, price-constrained queries and year-stamped queries trigger search at close to 100%.
- Discovery-shaped queries trigger at roughly seven times the rate of informational ones.
- Brand familiarity operates as a clean binary. An unknown brand forces a search; a familiar one is answered from training data. The same testing found “what is [brand]” ran no search while “tell me about [brand]” ran a deep one — a definition request signals confidence, an open-ended one signals uncertainty.
You are fetched where the engine is uncertain. Quality cannot enter the trigger decision: at trigger time the engine has not seen your page and does not know it exists. Everything the discipline optimises sits downstream of a gate closed on grounds having nothing to do with the document — and the behaviour is stable enough across the current generation of AI browsers and assistants to be worth designing around rather than waiting out.
Naming is a hedging device
A second mechanism explains why some true things get cited and others get absorbed. A model attributes what it will not assert in its own voice. Settled facts are folded in with no source; contested, specific, recent or numerically precise claims arrive with a name attached, because attribution is how a generator manages exposure. One analysis of Gemini-3 AI Overviews found 91% of answers correct while only 39% were fully supported by the sources cited — the citation was doing a rhetorical job as much as an evidential one.
Tracking data agrees from the other side: one monitoring study found brands are mentioned roughly three times as often as they are cited. That gap is reported as a measurement nuisance. It is what absorption looks like from outside — the engine knows the thing, so it no longer needs the source.
6. The consumption problem: you are cited during the interval of your own novelty
Publishing something genuinely new starts a clock. The claim is unpredictable, so retrieval fires and you are fetched and named. Being named is precisely what gets the claim restated — by outlets that cover it, competitors who rebut it, trade press that summarises it, forum threads arguing about the methodology. Every restatement makes the claim more predictable from the corpus. Eventually it is common knowledge: the model answers from its weights, in its own voice, with no name attached and no fetch fired.
Success is the mechanism of your own de-citation. A citation is not a reward for holding knowledge. It is a receipt for the engine’s ignorance, and the receipt expires when the ignorance does.
You are running a flow business, not building a library
The standard content-asset model treats published work as a stock that appreciates. A net-new claim does the opposite: maximum value at publication, decaying as it succeeds. The leading indicator of decay is the thing you were celebrating six months earlier, which is why teams modelling content as a stock over-invest in promoting claims already halfway absorbed. When a claim stops being named, the recovery playbook for lost AI citations is usually the wrong tool: nothing broke, the fact graduated.
What actually survives absorption
Two things. The first is a claim that cannot be restated without your data — a lane-level series, an incidence rate over a population only you observe, a figure that changes each period. Restating it requires the numbers, and the numbers arrive one period at a time, so absorption can only consume what you have already published. The second, and far more durable, is the position of being the party who gets asked. Absorption takes the fact. It does not take the meter, the sample frame, the consent agreement or the phone number. This is why the same ten organisations keep appearing in a sector’s evidence base decades after any particular number they published stopped mattering.
7. The Absorption Test, and the replacement rate it implies
The table tells you what to commission. This tells you how much, and how often. It takes an afternoon a month, and it is the only way to turn a claim about novelty into a budget line.
The Absorption Test
1. For each published claim, write the question it answers in the two phrasings that behave differently: a definitional one (“what is the excursion rate for chilled pharmaceutical transport in the UK”) and an open-ended one (“tell me about excursion rates in UK cold-chain transport”).
2. Run both across five engines in clean sessions, monthly, with personalisation and memory off.
3. Score every result into three states. Named: your claim, your name, ideally a link. Carried: your claim, asserted without you. Absent: the claim does not appear.
4. The naming half-life is the number of months until a majority of engines move a claim from Named to Carried.
5. Track Absent separately. It is not the opposite of Carried — it means the claim never propagated at all, a different failure and a cheaper one to fix.
Read the half-life against three bands. Above twelve months, the claim cannot be restated without your data: fund more of whatever produced it. Four to twelve months is normal for a fact anybody can carry once told. Below four months, you published something the corpus was about to produce anyway — a labour-row output priced as a standing-row one.
The arithmetic is the useful part. Suppose you want N claims named at any one time, each decaying with a half-life of h months. In steady state you need roughly 0.69 × N ÷ h new claims per month to hold the position. Twelve claims at a nine-month half-life is about one new claim a month; the same twelve at four months is about two. This is the number that quietly kills the annual flagship study — one report a year sustains roughly one claim, and no promotional budget alters the arithmetic, because promotion accelerates absorption rather than delaying it.
Stated limitation. The test measures naming, not truth, traffic or revenue, and any single month is confounded by engine-side changes you cannot see. Read the slope, not the level: a claim that drops out for one month and returns has told you about a model update; one sliding from Named to Carried across four consecutive months has told you about absorption.
8. Where this argument fails
The strongest objection is sharper than “novelty is unimportant”. If unpredictability is the operative variable, and unpredictability can be manufactured by computation, the winner is whoever runs the most measurements — and within two years that will be an automated pipeline, not a person. Row three is exactly what machines are good at: joining registers, assembling series, re-deriving figures at a scale no analyst can match. Concede it fully — computation-sourced novelty is already being industrialised, and “produce net-new material” will shortly be an instruction available to anyone with an API key.
Four things bound the objection, and none of them rescues the labour row.
- Standing is not compute-limited. You cannot compute your way into being the party the meter is attached to, or the party a hospital trust lets through the door. The green row is access-limited, and access is granted by people to people.
- The binding constraint is collection, not analysis. Scaling analysis while collection stays fixed produces more analyses of the same data — a race to the fat middle by a more expensive route.
- Volume collides with verification. The support gap above suggests engines retrieve a number and carry it imperfectly. At industrial volume nobody checks anything, and an unchecked figure is a claim rather than a finding.
- Being asked survives absorption. The residue identified above is a social position, conferred by other parties over time. No pipeline produces one.
The second objection is better. The retrieval corpus may degrade even if the weights do not. Training sets get curated by well-resourced teams; the live web fetched for grounding cannot be. If the fetchable web fills with machine-written material, retrieval quality can fall while model quality holds — an under-discussed asymmetry, and one reason labelling and disclosure regimes for AI content have attracted regulatory attention. That objection stands, but it does not rescue the human-premium thesis: a degraded retrieval layer does not read bylines either. What it raises is the value of the one thing surviving a degraded corpus — a claim with an identifiable owner and a checkable origin, which points at origin and custody infrastructure, not at byline strategy.
A limitation of my own. Predictability is assessed on outputs, not on the corpus. Showing that a model can produce a sentence close to yours tells you it was predictable; it cannot distinguish “the corpus already contained this” from “the model generalised well”. Treat a high score as a prompt to investigate, not a diagnosis.
9. What this changes about earned links
Two of the four sources of novelty require somebody who is not you, for different reasons.
Standing needs a third party because your own vantage is too small. You observe your customers, your fleet, your caseload. The population anybody cares about is the sector. An insurer’s claims file, a trade body’s survey, a regulator’s register, a university group’s instrumented sample — each is a vantage you cannot reach alone, and a joint study is the ordinary way to obtain one. What you buy is not a link but a share in an observation post, priced as capital expenditure rather than as a campaign, roughly how sponsorship-based link acquisition behaves when done properly.
Labour needs a third party because a computation nobody has checked is a claim. Publishing the method and inviting a re-run costs little and changes the class of the output. The cheapest version is a named methodology page and a standing offer to share the working file; the strongest is a second party who runs the operation independently and publishes what they got, differences included.
The prospecting screen this adds
Conventional screens — authority score, traffic, topical relevance — say nothing about whether an outlet holds or can obtain a vantage. Most of the standard link-building toolset cannot see it, because it is a property of the organisation rather than the domain. Add a column and ask one question of every prospect: does this party observe something, and can they be persuaded to publish what they observe? Trade bodies, insurers, professional registers, procurement functions and university research groups score badly on every conventional metric and hold most of the vantages in any sector.
It also demotes the formats the field defaults to. A contributed guest article and a placed link in an existing page both buy distribution for a claim you already own: useful, but no vantage and so no novelty. The effect compounds in multi-market work, because running one campaign across several countries multiplies distribution while leaving the underlying observation single-sourced.
Two rules fall out. Commission the series, not the study — an annual report has a naming half-life set by the calendar, an instrument that must be re-run has one set by the data. And aim at the query shapes that still trigger a fetch. Familiarity suppresses retrieval on definitional questions while comparison and constrained questions fetch almost every time, so the durable retrieval surface is comparative. Placements that put your figures inside somebody else’s comparison — a buyer’s guide, a specification note, a benchmark table — sit on the shapes that still open the gate. That is the underrated mechanism behind third-party listicle and roundup placements and behind the way assistants weigh product recommendations.
10. Worked example: Trevanion Cold Chain
Trevanion Cold Chain is a temperature-controlled logistics operator in Bridgwater, Somerset: £16.8M turnover, 61 refrigerated vehicles, a GDP-licensed pharmaceutical lane alongside chilled food. Every consignment carries a calibrated data logger and every excursion is documented by obligation.
The premise they bought first
From September 2025 they ran the human-premium play: sixty articles over seven months, bylined by named depot managers, £41,000 all in, on the explicit theory that authentic human writing would be repriced as the web filled with machine text. By April 2026 they were named in 3 of a 40-prompt tracking set. A predictability check across their forty strongest claims found 31 could be reproduced almost exactly by a model working from public material alone. They had spent £41,000 writing the corpus back at itself in a better voice.
What they actually held
Four years of calibrated logger telemetry across 38,412 monitored consignments: excursion incidence by route, season, trailer type and dwell point. Nothing in the public record could predict any of it. In May 2026 they published the Lane Excursion Report, with the excursion definition on its own page and the breakdown at route level. Naming moved from 3 to 19 of 40 prompts in nine weeks, and the most reproduced line was the winter M5 corridor figure, which no trade publication had ever had access to.
What the Absorption Test showed three months later
By August 2026 the headline national figure was Carried rather than Named by 3 of 5 engines — a naming half-life of roughly seven months. The lane-level figures were still Named by 5 of 5, because restating them requires the series, and the series exists only inside Trevanion’s telemetry. Twelve named claims at a seven-month half-life needs about 1.2 new claims a month, so they moved to a monthly corridor bulletin at £3,100 and commissioned two co-held studies: an insurer’s claims file covering spoilage across 400 operators, and a university logistics group re-running the corridor figure against independently instrumented vehicles.
What went wrong
- Absorption upward was faster than absorption into a model. A trade association restated the headline figure in a policy submission without attribution, and within four months the association’s version was the reference every journalist quoted. Absorption into an institution proved quicker and more total than absorption into a set of weights, with no recourse.
- The most defensible data was the most restricted, and the restriction sat with the suppliers. Two pharmaceutical customers objected to lane-level publication on commercial-sensitivity grounds and the quality agreements gave them a veto. The anonymisation required to satisfy them destroyed the route granularity — the exact property that made the figure unabsorbable.
- The replacement rate outran the genuine discovery rate. Months three and four contained no real signal and were padded to fill the slot. Neither was cited, and the two dragged the series’ naming rate for a quarter. A replacement rate is a ceiling on what absorption will let you hold, not a target to hit regardless of whether anything happened.
- The labour row decayed five times faster than the standing row, on an identical budget. A competitor replicated the public-weather join in seven weeks from public inputs and published a broader version. That novelty lasted two months; the telemetry figures were still Named nine months on. Both were commissioned at the same price, on the same brief, in the same quarter — nobody had drawn the distinction when commissioning.
11. What to do on Monday
- Take your twenty strongest claims and run the predictability check: strip the payload from each headline sentence and ask a model to complete it from public material. Anything it completes accurately, the corpus already contained.
- Sort every content line item in the current plan into the four rows and total the spend by row. Most plans turn out to be mostly authorship and time.
- List every population you observe as an operational by-product — meters, tickets, inspections, claims, deliveries — and check what is retained and for how long. Retention is usually the binding constraint, fixed by a policy nobody has revisited.
- Set up the Absorption Test this month: forty questions, two phrasings, five engines, three states. It is the baseline you will need in six months, and it cannot be reconstructed retrospectively.
- Compute your replacement rate from the first half-lives you observe and compare it with what you are commissioning. The gap is your future de-citation.
- Convert one planned annual study into a recurring instrument with a fixed definition page, a stated cadence and the method published alongside it. Then re-score your prospect list on the vantage question and expect the ranking to invert.
- Open a conversation with one third party holding a vantage you cannot reach — a co-held study, not a placement. The fundamentals of earning links and mentions still apply, and so does the reality that this is a specialist function rather than a channel task.
- Check what your site does when an assistant arrives and cannot resolve a figure: agentic browsing without a clear answer produces a fetch with no citation, the worst of both positions.
The collapse literature is interesting and almost entirely irrelevant to the decision in front of you. What matters is simpler: an engine fetches what it cannot predict, names what it will not assert, and stops doing both the moment your contribution has done its job. The only durable position is to be the party who keeps producing the next unpredictable thing — and, more than that, the party everyone else has to ask. That is a decision about what your organisation measures, retains and is willing to publish, and it looks far more like local and sectoral corroboration than like anything in a content calendar.
