AI Slop Citation

Citation in the Age of AI Slop: Standing Out When Everyone Publishes at Scale

TL;DR

The flood arrived and the citation layer held. Roughly half of all new articles are now machine-written, yet 86% of the articles ranking in Google and 82% of those cited by ChatGPT and Perplexity are human-written.

The comfortable explanation is false. Every headline study behind that finding used a style detector. The measured penalty therefore falls on a style, not on a provenance, and no engine has ever claimed to detect the latter.

The pattern is a tie-break, not a filter. Near-parity in the top ten, an eight-fold gap at position one, the widest gap of all at the citation layer. A filter removes early and uniformly. This one only bites where slots are scarce.

So the threat is resemblance, not competition. Slop converted previously neutral structural features into markers of a discard class, and the optimisation-checklist genre manufactures exactly that commonality.

The link consequence is mechanical. Abundance devalued what anyone can produce and revalued what someone else must concede, which repriced earned placements upward without a single algorithm change.

The flood arrived. The citation layer held.

In 2022 Europol forecast that 90% of online content would be synthetically generated by 2026. It is 2026, and the forecast was wrong in an interesting way: not because the machines underdelivered, but because the number everyone quoted described production while the number that matters describes selection.

Graphite classified 65,000 English-language URLs drawn from Common Crawl and published between January 2020 and May 2025. The machine-written share climbed from roughly a tenth in late 2022, briefly overtook human writing in November 2024, and then stopped. It has hovered near even for five straight quarters. The wave broke and flattened.

The second finding is the one that reorganises the field. Across ranking and answer surfaces, 86% of the articles appearing in Google results were human-written, and 82% of the articles cited by ChatGPT and Perplexity were human-written. When machine-written articles do rank, they rank lower. So the supply of published text became roughly half synthetic while the supply of selected text stayed overwhelmingly human. Something in between is sorting, and the entire practical question is what.

What is AI slop?

AI slop is high-volume, low-marginal-effort published text whose economics only make sense at scale. It is not defined by the tool used to write it. It is defined by the absence of anything that would have cost the publisher something to produce, which is why a hand-written page produced to a template can be slop and a heavily edited machine draft need not be.

That distinction is not pedantry. It determines whether the correct response is to write differently, publish less, or buy something you cannot write at all. Most teams have assumed the first, and the evidence points hard at the third.

Every headline study measured a style, not a provenance

Before building a strategy on the 86% figure, look at how it was produced. Graphite ran a detector over its corpus and applied a 50% threshold: above it, the article was counted as machine-written. Semrush ran GPTZero across 42,000 blog pages covering 20,000 keywords. Originality.ai scanned top-20 results and reported around 17% synthetic at the time of sampling.

Every one of those instruments is a classifier trained on surface features: sentence-length variance, connective density, lexical predictability, the absence of digression. None of them observes how a document was made. They observe how it reads. A study using such a tool cannot separate a page written by a model from a page written by a person imitating the register that models produce, because at the level of the input those two artefacts are the same object.

Graphite says so directly. Its methodology notes that it did not evaluate heavily human-edited machine drafts and that such content may well perform. Semrush’s write-up concedes the same point in one sentence: the detector does not care about your process, it reads the finished product. That is an admission the field has quoted past for a year. It means the widely repeated claim that engines demote AI content is unsupported by the very studies used to support it.

Ahrefs killed the binary from the other end. Across 600,000 pages it found that 86.5% contain some machine-generated text, while only 4.6% are wholly machine-written. Almost the entire web is now a hybrid. A methodology that sorts hybrids into two bins by a threshold is not measuring provenance at all; it is measuring where a page sits on a style gradient, and then reporting the result as though it were a fact about authorship. Anyone reading these figures alongside the wider 2026 link building statistics should apply the same discount to every AI-content share quoted anywhere.

Key takeaway

The evidence base supports one claim only: pages that read like unedited model output are selected less often. It supports no claim whatsoever about pages that were model-drafted. Those are different populations, and your content probably sits in the gap between them.

The pattern is a tie-break, not a filter

Three measurements exist. Nobody has lined them up, and lined up they say something none of them says alone.

In the top ten, near-parity. An earlier Semrush study found around 8% of sampled URLs were likely machine-written, and that 57% of that content appeared in the top ten against 58% of human content. A one-point difference is nothing.

At position one, an eight-fold gap. In the larger 42,000-page study, GPTZero classified 80% of position-one pages as human and 9% as purely machine-written.

At the citation layer, the widest gap of all. 82% of cited articles human-written, against a published population that is roughly half machine-written.

What a filter would look like in this data

If engines were detecting and demoting synthetic text, the effect would appear early and uniformly. Pages would be suppressed at the point of classification, and the suppression would be visible everywhere at once: in the top hundred, the top ten and position one alike, in roughly the same proportion. A filter is a property of the document. It does not know how many competitors are queued behind it.

What a tie-break looks like instead

What the data actually shows is a penalty that scales with scarcity. Where there are ten slots, the difference is a rounding error. Where there is one slot, it is eight-fold. Where an answer carries between two and seven citation slots, it is wider still. That is the signature of a decision made after the text has stopped discriminating: when several candidates are adequate, something other than the text breaks the tie, and machine-register pages lose that tie-break systematically.

This reframes the whole problem. You are not being screened out for what you wrote. You are being sorted last among things that already qualified. The distinction matters because the two failures have opposite remedies: a filter is escaped by writing better, and a tie-break is escaped by carrying something the other candidates do not. It is the same asymmetry that governs how AI Mode assembles a diverse citation set — relevance qualifies you, and something else selects you.

Screening economics: why the sort exists and which way it fails

A screen is not chosen because it is right. A screen is chosen because it is cheap, and because it fails in the direction its owner can afford. Every triage system in commercial use follows that rule, and retrieval is no exception.

The binding constraint here is not money, which is the assumption most commentary makes. It is latency. A user waiting for an answer imposes a hard time budget, and retrieval-augmented generation — fetching live pages at answer time and reading them before responding — has to fit candidate selection, fetching and synthesis inside that budget. Google’s own May 2026 guidance on generative search states that AI Overviews and AI Mode are rooted in core Search ranking and quality systems, use retrieval-augmented generation with query fan-out, and require no special markup or llms.txt file. Read that as an architecture statement rather than a reassurance: whatever the core index has already decided about your page is inherited by the answer layer before any model reads a word of it.

A pre-filter that cheap will produce false negatives, and here is the uncomfortable part: those false negatives cost the engine almost nothing. A missed good page is invisible to the user, who never learns what was not shown. A published bad page is highly visible. Any operator facing that asymmetry will tune the screen to over-discard, because over-discarding is the failure nobody can see. Yours is one of the pages nobody can see.

Does Google penalise AI content?

No, and it has been explicit about this since 2023. The scaled content abuse policy introduced in March 2024 and effective from 5 May 2024 is deliberately method-agnostic — it targets content produced at scale to game rankings, no matter how it is created. The January 2025 rater guidelines assign the lowest quality rating to pages that are completely or nearly completely automated without effort, originality or added value, which is a judgment about effort, not about tooling.

The enforcement is real even though the policy is not about authorship. Google stated a target of cutting low-quality, unoriginal results by around 40%, later revised upward to 45%; manual actions in this class deindex entire domains, and in the March 2024 wave 837 of 49,345 monitored sites were removed outright. The March 2026 core and spam update made scaled content abuse the primary enforcement target. None of that requires a provenance detector, which is precisely why none of it can be escaped by writing more humanly.

Your real threat is resemblance, not competition

Competition is a story about rivals taking your slot. Resemblance is a story about being mistaken for something else, and it behaves nothing like competition — you can lose to resemblance in a category with no rivals at all.

Slop imposed a negative externality that almost nobody has priced. Before the flood, a page carrying a keyword-led H2 structure, a summary box, a comparison table and a tidy conclusion carried no adverse signal, because those features were expensive enough that only deliberate publishers produced them. Once they became the default output of a generation pipeline, they stopped being neutral. They became statistically associated with the discard class. Nothing about your page changed. The reference distribution changed underneath it.

The bitter irony is that the optimisation genre now manufactures precisely the commonality that hurts. A widely followed standard is a machine for producing uniformity, and a uniform is what a cheap classifier is best at recognising. Every checklist promising a format that generative engines prefer is, in aggregate, an instruction to converge — and convergence is the input a discard heuristic is tuned on. This is why teams following current guidance with real discipline sometimes find their most carefully engineered technical foundations performing worse than a plainer estate that ignored the advice.

Which raises the operational question this article exists to answer. You cannot see the discard heuristic, and you cannot ask about it. But you can sample its output.

Instrument 1: The Discard Sample

In 1943 the Statistical Research Group was asked where to add armour to bombers, using damage maps from aircraft that returned. Abraham Wald pointed out that the maps described survivors: the undamaged regions were undamaged in the sample precisely because hits there did not come home. Armour belonged where the returning planes showed no damage at all.

Search marketing has been reading returning bombers for twenty years. Every ranking study, every SERP teardown, every competitor backlink analysis describes pages that survived a selection process, and infers the rule from the survivors. That method cannot work here, because the question is what gets discarded. Structurally, competitor analysis cannot answer it. You have to sample the dead.

THE DISCARD SAMPLE — a five-step procedure

1. Fix the query set. Choose 40 to 60 queries you would expect to be cited on. Freeze them in writing before you look at anything, because a query set chosen after the fact will quietly drift toward the ones you already win.

2. Build the survivor set. Run each query across the engines your buyers actually use and record every page named or cited. This is the set the whole field already studies.

3. Build the discard set — the step nobody does. For each query, collect the pages that were plainly topically eligible and were not named: pages ranking on page one or two of conventional results for the same query, pages from the same publishers whose other work does get cited, and pages your own team would have shortlisted as answers. These are the aircraft that did not come home.

4. Extract the structural markers. Compare the two sets on form rather than content: heading patterns, paragraph length distribution, table and summary-box presence, publication cadence, author attribution, whether any figure in the page originates outside the domain. Keep every marker present in at least 60% of the discard set and under 30% of the survivor set. Ten to fifteen markers is typical.

5. Score your own last twenty pages. Your RESEMBLANCE SCORE is markers carried divided by markers identified. Under 20% is clear. Between 20% and 40% is a watching brief. Above 40% means your production standard is optimising you into the discard class, and no amount of additional care inside that standard will help.

One caveat, stated plainly because the instrument is easy to over-read: the Resemblance Score is correlational. A high score tells you your pages look like the discarded ones on features that may or may not be causal. It is an investigation trigger, not a diagnosis, and its main value is that it points at form when every other tool in the stack points at content.

Instrument 2: The Discard Order

A page can fail at five distinct points, and the field routinely treats them as one event called not ranking. They are not one event. They differ in what the system could see at the moment it decided, and — decisively — in whether better writing can move them at all.

StageWhat the system can see at that momentCan better writing move it?Where you can see the failure
1. Not crawledA URL and whatever links point at it. No content at all.No — nothing has been readCrawl stats; server logs
2. Crawled, not indexedThe full page, judged against everything already stored on the topic.No — the judgment is comparative, not qualitativeSearch Console index report
3. Indexed, never retrievedAn embedding and a set of ranking signals. The prose is not re-read.Rarely — retrieval acts on stored representationsNowhere directly; only by absence
4. Retrieved, not selectedThe passage, alongside two to seven rival passages competing for the same slot.Yes — this is the tie-breakAnswer monitoring; cited-set tracking
5. Selected, not namedYour content, absorbed into the answer without attribution.Yes — naming responds to how a claim is writtenAnswer text; unlinked mention tracking

The punchline is in the third column. Craft acts on the two green rows and nowhere else. Everything above them is decided before anyone reads your sentences, which is why so much editorial effort produces no observable movement — it is being applied to a stage that does not consume it.

Row two deserves particular attention, because it is where most large estates die and it is routinely misread as a technical fault. First-party analysis by Adam Gent across 1.4 million pages on 18 sites found that a page not recrawled within 190 days has roughly a 90% chance of being forgotten entirely, and that 70 to 80% of URLs sitting in the crawled, currently not indexed state had previously been indexed and were actively removed. He suggests the status would be better named crawled, previously indexed. Gary Illyes has confirmed that Google purges low-value URLs and can forget it ever saw them.

Key takeaway

An index discard now cascades into the answer layer, because the generative surfaces sit on the same core systems. A page that has been forgotten is not competing badly for citations. It is not a candidate.

One discard, three indexes

A discard is index-specific, and that is operationally useful. ChatGPT retrieves against Bing, Claude against Brave, and Google’s answer surfaces against Google’s own index. Each runs its own retention policy over its own crawl budget, so a page forgotten by one is not necessarily forgotten by the others. Two consequences follow. A resemblance problem will present unevenly across engines, which means a single-engine check will mislead you about its size in both directions. And the cheapest early recovery route is usually whichever index still holds you: if Brave has retained pages Google has dropped, corroboration built now lands somewhere immediately while the slower recovery runs in the background. Check all three before concluding that a page is dead.

What abundance did to the price of an earned placement

There is a clean economic reading of all of this, and it does not depend on any prediction about model behaviour.

A signal carries information in proportion to what it costs to produce. Text on your own domain used to cost something: time, expertise, editorial attention. It now costs almost nothing, so its informativeness collapsed — not because engines decided to trust it less, but because it stopped separating anyone from anyone. Meanwhile the cost of a third party publishing your figure, naming your firm or reproducing your method did not fall at all. It still requires a decision by someone with their own reputation at stake. Understanding what backlinks actually are as concessions rather than votes makes the consequence obvious.

So earned placements were repriced upward mechanically, by the collapse of the substitute rather than by any change in their own properties. Abundance devalues what can be produced and revalues what must be conceded. Muck Rack’s analysis of more than 25 million cited links across ChatGPT, Claude and Gemini in 17 industries found earned media accounting for 84% of AI citations, with journalism alone at 27%. Moz’s 2026 work adds that 88% of AI Mode citations do not come from the organic top ten — which means the answer layer is reaching past the ranking layer to find corroboration, and the relationship between AI Overviews and backlinks is stronger than the ranking data alone would suggest.

There is a catch, and it is expensive. If pages can be discarded before they are read, then some of the placements you are buying sit on pages that no longer exist as far as the answer layer is concerned. Nobody in outreach checks for this. It should be the first check, not the last.

THE CARRIER CHECK — four questions before paying for any placement

1. Is the exact host page indexed? Not the domain — the page. Query the specific URL. A live page on a healthy domain can already be in the forgotten state, and a link on it is a link on an artefact that is not a candidate for anything.

2. Is the outlet’s recent output indexed at roughly 70% or better? Sample twenty URLs published in the last ninety days. An outlet whose new work is not being retained is running down its own eligibility, and your placement will age into that.

3. Is the outlet ever named in an answer on its own core subject? Ask the engines ten questions the outlet plainly covers. If it is never named on its home turf, it will not carry you anywhere on yours.

4. What is posts-per-week divided by named editorial staff? This is the cheapest slop proxy available and it is public. A ratio that no group of named humans could plausibly sustain tells you what the outlet is, and increasingly what the classifiers think it is.

You are not buying a link. You are buying a page’s continued eligibility to be read.

Applied honestly, this check reorders a prospect list faster than any authority metric, and it is unusual in being uncorrelated with everything the field already screens for. Two outlets with identical scores in every tool listed in a standard link building tools comparison can sit on opposite sides of it.

Cadence is a discard marker, so publishing more makes it worse

The reflex response to a visibility problem is volume. Under the mechanism described here, volume is the one lever that reliably deepens the injury, because publication rate is itself one of the structural markers a cheap screen can read without opening a single page. It is available in a sitemap. It costs nothing to compute. It correlates with exactly the class the screen exists to remove.

A defensible rate is the rate at which named people can genuinely read what goes out. That is not a moral position, it is a signalling one: reviewability is the property that is expensive to fake, and it happens to be the property that leaves public traces. It also changes what you spend on. When the estate stops growing, budget moves toward assets whose value does not decay with volume — the interactive calculators that earn links at scale, the scrollytelling formats that a template cannot mass-produce, and the reactive work in newsjacking where being first is the whole asset.

Note also what this does to link acquisition rhythm. A sudden burst of placements against a static estate reads very differently from steady accumulation; the considerations in link velocity for 2026 apply with more force when the estate has stopped absorbing new pages.

Worked example: Ormskirk Control Systems

Ormskirk Control Systems builds industrial control panels and machine-safety systems from Skelmersdale. Turnover £12.6M, 74 staff, roughly 60% of new enquiries historically sourced online. Through 2025 it followed the standard advice and scaled its estate from 210 pages to 1,900.

By January 2026, 1,347 of those 1,900 pages sat in crawled, currently not indexed. That alone would have been a familiar story. The finding that changed the diagnosis was this: 148 of the 210 original engineer-written pages were in there too. The human pages died alongside the machine pages.

The explanation was a 2024 template retro-fit. When the estate was rebuilt, the older engineer pages were reformatted into the same house structure as everything that followed — same heading skeleton, same summary block, same comparison table, same closing section. The engineering content survived. The uniform was applied over it. A Discard Sample across 48 queries returned a Resemblance Score of 61% across the estate and 56% on the engineer-written pages specifically, which is the number that tells you the problem is form and not substance.

The Carrier Check produced a second, costlier finding. Of 61 placements bought over two years at a cost of £38,000, 19 sat on pages that were not indexed at the time of checking. Thirty-one per cent of the link budget had purchased URLs rather than links.

What they did, and what it cost

The estate was cut from 1,900 pages to 340. Publication stopped entirely for eleven weeks. £54,000 went into three artefacts that a third party had to produce, on the reasoning that nothing they could write themselves would break a tie-break: a failure-mode dataset built with their insurer across 1,180 inspected panels, a standards note co-authored with a trade body, and an arc-flash containment figure re-run independently by two test houses.

Over the following two quarters, naming across a fixed 40-prompt set moved from 4 to 22. The single most reproduced item was not the dataset headline but the independently computed containment figure — the number they had the least control over.

The four things that went wrong

First, the cull deleted 88 pages that between them held 41% of the site’s referring domains, because no backlink audit was run before the delete list was drawn. Thirty-four of those domains proved unrecoverable. Second, the eleven-week freeze cost a seasonal term the company had held for three years. Third, index recovery preceded citation recovery by about five months, producing two quarters in which the dashboards looked repaired and the outcome did not — enough to generate real board scepticism at exactly the wrong moment. Eligibility is not retrieval. Fourth, the insurer dataset carried a twelve-month exclusivity clause that blocked an update, so the headline figure went stale and a competitor’s newer number displaced it in several answers.

Key takeaway

The recovery worked because it replaced text the company could produce with facts other parties had to concede. The failures were all failures of sequencing: audit before you delete, and never sign an exclusivity that outlives the freshness of the number it covers.

Where this argument breaks

The strongest objection is not that the screen does not exist. It is that the screen is temporary. Inference costs have fallen by orders of magnitude and continue to fall; when reading every candidate properly becomes cheap, the crude structural pre-filter dissolves, and with it the entire resemblance problem. On that view this article describes a transitional artefact and the correct strategy is to wait.

It is a serious objection and it is partly right. Four things bound it.

Latency, not cost, is binding. A user waiting on an answer sets a time budget that cheaper tokens do not relax. Reading a thousand candidates carefully is a wall-clock problem before it is a budget problem, and wall-clock is what the product is competing on.

Reading cannot verify. Oumi’s audit found that 91% of Gemini-3 AI Overviews contained the right answer while only 39% were fully supported by the sources cited. More careful reading closes the gap between the source and the summary; it does nothing about whether the claim is true, which still requires a party outside the document. Corroboration is not a rendering problem.

The share would already be drifting. If cheap reading were dissolving the screen, the human share of citations should be sliding toward the roughly 50% share of the published population. It has not moved off 82%. Something is holding, and the most economical explanation is that the selection criterion was never really about text.

Regulation pushes the other way. The UK Competition and Markets Authority’s fair-ranking commitments and the EU’s transparency duties both require selection to be explicable. Explicability favours legible, auditable, non-text criteria — publisher identity, licensing status, corroboration count — over an unreproducible judgment about prose quality. The compliance direction entrenches the proxy rather than dissolving it.

Where the objection wins: if you have genuinely distinctive substance trapped inside a conventional format, the resemblance penalty on it will decay over time. That is an argument for fixing form now and not rebuilding your entire content strategy around a screen that may soften. It is not an argument for waiting, because the link-side consequence — the repricing of conceded signals — follows from abundance itself and survives the screen entirely.

What to do on Monday

A literal sequence. The first three cost nothing but time.

  • Pull the index report first. Count how many of your published URLs sit in crawled, currently not indexed, and check specifically whether your best hand-written pages are among them. If they are, you have a form problem, not a quality problem.
  • Run the Discard Sample on 40 to 60 frozen queries. Build the discard set before you build the survivor set, so the survivors do not anchor your reading of what matters.
  • Score your last twenty pages for resemblance. Above 40%, stop the template before you commission anything else.
  • Run the Carrier Check across every placement bought in the last twelve months. Report the proportion of your link budget that landed on unindexed pages. That single number usually funds the rest of the work.
  • Audit referring domains before any content cull. Never delete a page whose backlink profile has not been checked and redirected. This is the mistake that is genuinely unrecoverable.
  • Reset cadence to reviewability. Set the publication rate to what named reviewers can actually read, and put the names on the pages.
  • Move the freed budget to conceded signals. Commission one artefact this quarter that a third party has to produce, measure, or restate — and refuse any exclusivity clause longer than the useful life of the figure.
  • Track naming, not traffic. Fix a prompt set and count how often you are named. Traffic will be the last thing to move, and by the time it does the decision that mattered will be five months old.

The wider point survives every uncertainty in the mechanism. When production becomes free, production stops being evidence. The only signals left with any information in them are the ones somebody else had to give up something to make — which is what link building was always doing, and why the strategies that earn rather than manufacture are the ones that came through the flood intact.

For the adjacent mechanics: how absence is located across surfaces is covered in the SERP-less audit; how carry works across a conversation in multi-turn query chains; how research modes assemble their sources in AI deep research citations; what happens to demand modelling once clicks stop arriving in the zero-click traffic model; where verification genuinely does and does not pay in signed versus unsigned content; and why a community that polices its own contributions still works, in Hacker News link building.

Leave a Reply

Your email address will not be published. Required fields are marked *

Trust Graph Previous post The Trust Graph: How Engines Will Weight Verified Sources in 2027