Citation Poisoning

Citation Poisoning: How Bad Actors Manipulate AI Answers — and How to Detect It

TL;DR

Citation poisoning is content published to be retrieved and repeated by an AI answer engine. It attacks not what a model believes but the eight or ten documents it reads before answering one question.

The standard response — monitor, find the bad source, correct it — begins with a step that fails about half the time: in Lasso Security’s June 2026 experiment, 48.2% of runs carrying a planted false claim never cited the page it came from.

Exposure is measured in two numbers, not in sentiment: how much of the evidence behind your commercial answers a stranger can edit today, and how many independent documents assert each claim about you.

The defence is coverage, not correction — which is the same asset you were already buying for demand.

The loop that begins with a broken step

The standard 2026 advice for a brand that finds itself misrepresented inside an AI answer runs in four steps. Track your brand across ChatGPT, Perplexity, Google’s AI Overviews and AI Mode with a visibility tool. When a false claim appears, open the citations and find the source. Contact that source, or publish a correction that outranks it. File feedback with the engine. It is sold by every monitoring vendor in the category, and step two — the step everything else depends on — does not reliably work.

In June 2026, Lasso Security Research published the largest controlled test of the question to date. The team planted one false medical claim in an ordinary-looking cookery site and measured how often five production models repeated it when the page was mixed into a batch of real retrieved sources. Across 5,525 runs, 1,010 produced a detectable signal. Of those, 48.0% cited the planted page without repeating the claim, 3.8% did both — and 48.2% repeated the claim while citing something else entirely. In 487 runs the model promoted the claim with the page sitting in its context window and never named it once, attributing the idea instead to the supporting articles the researchers had planted around it, or to nothing at all.

Roughly half the time a planted claim travels, the citation list points at somebody else. The origin is not in the footnotes. This is not a bug that will be tuned away: retrieval-augmented generation, the technique of grounding an answer in documents fetched at question time, attributes loosely, so the sentence you are tracing and the link the model shows you are only sometimes the same object.

The most expensive live demonstration of this belongs to a solar installer in Isanti, Minnesota. Wolf River Electric sued Google in March 2025 over an AI Overview stating the company faced a lawsuit from the Minnesota Attorney General. The company had never been a defendant; the state’s actual case named four solar-lending firms. The complaint’s central factual allegation is the part that matters here: the AI Overview cited four sources, and none of them contained the claim. Wolf River’s initial disclosures put damages between $110 million and $210 million. There was no poisoned page to find, no publisher to email, no correction to request. The loop had no step two.

You cannot diagnose a damaging AI answer from the answer. Detection has to run in the other direction — from the question, into the set of documents that could be retrieved for it, and only then to the output. It changes what you buy: not a dashboard that watches sentiment, but an inventory of what exists to be read about you. Tools that watch for competitor misinformation in AI answers still have a job — they tell you that something is wrong. They cannot tell you why, and the industry has been selling the second thing while shipping the first.

What citation poisoning actually is

What is citation poisoning?

Citation poisoning is the deliberate publication of content built to be retrieved and repeated by AI answer engines, in order to change what those engines say about a person, product or company. It differs from ordinary misinformation in its target: the audience is the retrieval layer, not a human reader, and success is measured in citations rather than traffic.

When an engine answers a commercial question it does not consult the whole web. It runs one or more searches, pulls back a handful of candidate documents, and writes from those. That candidate set is small — often eight to twelve passages — and it is query-local: it is assembled fresh for that question, from whatever ranks or embeds closest to it. An attacker therefore never has to out-publish the internet. They have to out-publish the shortlist for one question.

Academic work put a number on how small that job is. In PoisonedRAG, presented at USENIX Security 2025, Wei Zou and colleagues injected five malicious texts per target question into knowledge bases holding millions of clean documents — 2,681,468 in the Natural Questions corpus, 8,841,823 in MS MARCO — and drove the model to the attacker’s chosen answer in up to 97% of cases in the black-box setting. Five against 2.7 million. Corpus size is not protection, because corpus size is not what the retriever is choosing from.

Two conditions have to hold for the attack to land, and both are worth naming because your defences attach to different ones. The document has to be retrieved, which is a ranking and similarity problem, and it then has to be believed, which is a corroboration problem. Their baseline run — the false page injected into the retrieved set with no optimisation applied — produced 0% citation and 0% claim repetition across all five models tested. The lie alone did nothing. What moved the models was presentation and agreement: reformatting the same claim from a buried paragraph into the first position of a ranked list flipped one model from ignoring it to recommending it, with the same query and the same ten sources.

The strongest single technique in that study was not technical at all. The researchers called it Editorial Endorsement: publish three short articles in the voice of plausible niche publications, each citing the target page and repeating the claim. That alone produced the claim in 92% of runs on Llama-4-Maverick (46 of 50) and 78% on GPT-4o-mini (39 of 50), rising to 98% and 84% when combined with a listicle layout. The cost, in Lasso’s own summary: a website and three fabricated editorial pages on plausible domains, a day’s work for one content marketer.

Key takeaway

An attacker does not need volume, authority or access. They need to be in the shortlist for one question, and they need something that looks like agreement once they are there. Both are cheap precisely where you are thinly covered.

The source walk: four causes, four remedies, one window

Because the citation list is unreliable, the first job after a damaging answer surfaces is a differential diagnosis rather than a takedown: the source walk. You take the specific damaging sentence, run the question that produced it across engines several times, collect every source cited, and then search independently for the claim itself — not for your brand. What you find sorts into four causes. They look identical in the answer and need completely different responses, and each has a wrong remedy that feels productive while burning the fortnight in which the claim is spreading.

CauseWhat the source walk findsWhat actually fixes itThe move that wastes the window
Synthesis errorThe claim appears in none of the cited sources, and nowhere else you can find. No author owns it.Publish a specific, dated, retrievable counter-fact; report through the engine’s feedback path; document the loss.Hunting for a source to contact. There isn’t one, and you lose a fortnight looking.
Ambient errorA real third party got it wrong in good faith — a trade write-up, a directory record, a stale acquisition note.Correct at the publisher. One email to a named editor removes the document from every future retrieval.Publishing your own rebuttal first, while the wrong page stays live and keeps being read.
Silence fillNo false source at all. The engine invented a specific figure because you publish none and the question demands one.Publish the number, with definition, period and date. The gap closes because the gap was the vulnerability.Treating it as an attack. There is no attacker, and no takedown will ever succeed.
Hostile plantSeveral thin pages, recently created, oddly shaped like the query, citing each other and nothing older.Platform enforcement against the host, on the host’s own compliance exposure — not persuasion of the author.Arguing with the poster, or issuing a public denial that gives the claim a better-ranking home.

The source walk. The same damaging sentence has four possible origins; only one has an attacker, and only one is fixed by a takedown.

Synthesis error is the common case, not the rare one

Most brands who believe they have been attacked have been mangled instead. BrightEdge analysed hundreds of millions of prompts between mid-January and February 2026 and found Google’s AI Overviews 44% more likely than ChatGPT to express negative sentiment about a brand, with brand controversies and legal issues the leading trigger at 32% of negative mentions across both platforms. Budget accordingly: the highest-frequency cause of a damaging answer is not a competitor.

The hostile plant has tells, and they are structural

A genuine plant is recognisable from its shape, not its content. The pages are young, and they arrive in clusters within days of each other. They are written in the exact grammar of the target question rather than of a subject, because that is what gets them retrieved. They reference each other and nothing older than themselves. They carry no independent trace: no companies-house record, no author with prior work, no earlier version in the Internet Archive. This is the same forensic instinct that spam-link detection has used for years, pointed at a different artefact — and the same instinct behind watching unnatural velocity in a backlink profile, where the giveaway was never the individual link but the arrival pattern.

One market wrinkle makes the tells harder. Paid placements on genuine third-party pages — the niche-edit trade — let an attacker buy a sentence inside an established article with real age and real authors, defeating every youth test above. What it cannot fake is the claim’s history: a fabricated fact about your certification has no earlier existence anywhere, however old the page carrying it. Walk the claim, not the domain.

In the UK, the fastest lever is the platform’s own liability

When a plant lands on a review platform — which is where plants land, because reviews are trusted, indexed and open — British law gives you a lever that most GEO advice misses entirely. Under the Digital Markets, Competition and Consumers Act 2024, in force since 6 April 2025, publishing fake reviews is a banned practice, and the ban explicitly covers reviews that are not based on genuine experience whether positive or negative. The Act also places a positive obligation on platforms hosting reviews to take reasonable and proportionate steps to prevent, detect and remove them. The Competition and Markets Authority can fine up to 10% of global turnover through administrative proceedings without going to court, and it is using it: five investigations opened on 27 March 2026.

The operational point is about who you write to. A complaint arguing that a review is unfair invites a judgement call from a moderator. A complaint that identifies a fabricated review, evidences that no transaction exists, and frames removal as the platform’s own duty under a statute that carries turnover-percentage fines invites a compliance decision instead. It is the closest thing in this field to the leverage a manual action recovery gives you with Google: not an appeal to fairness, but a documented case that the other party has a problem.

The exposure metric: your writable share

What is a writable source?

A writable source is any document behind your AI answers that a stranger can add to or alter today without your permission — a forum thread, a review profile, a wiki entry, a wire release. Your writable share is the percentage of sources cited across your commercial questions that fall into that class, and it is the single most useful number in this entire discipline, because it is the fraction of your reputation that is available for editing by people who do not work for you.

The published estimates disagree so violently that quoting any of them at your board would be malpractice. Peec AI’s analysis of 30 million cited sources ranks Reddit the single most-cited domain across every major engine. Evertune’s study of 200 million prompts finds no domain exceeding roughly 5% of citations, with Wikipedia, Reddit, LinkedIn and YouTube combined rarely topping 5%. Yext’s analysis of 6.8 million citations puts forums at 2% and brand-controlled sources at 86%. Which is exactly why the number has to be yours, computed on the forty questions your buyers actually ask, and recomputed quarterly. Any figure you take from a published statistics round-up is a fact about somebody else’s market.

Source classTime for a stranger to publishReview gateWho can get it removed
Community forum threadMinutesNone; volunteer moderation after the factThe moderator, at their discretion — and the post outlives the account
Review platform profileMinutes to hoursAutomated fraud checks, patchyThe platform, fastest when framed as its own statutory duty
Wire or syndicated releaseSame day, for a feeEditorial checks vary by wire; none re-check the syndicated copiesNobody, in practice — the copies outlive the original
Collaborative encyclopaediaMinutes to edit, days to surviveCommunity reversion, sourcing rulesEditors, quickly, if the claim lacks a citable source
Independent trade publicationWeeks, and only with a real storyNamed editor with a byline to protectThe publisher, on correction — and they are contactable

The classes at the top of this table are what move your writable share; the class at the bottom is what lowers it.

Two features of writable sources make them worse than their share suggests. They are durable: Profound’s data shows the Reddit posts engines cite skew old, averaging around a year, with 4% dating from 2019 or earlier — the notorious AI Overview recommending glue on pizza traced to a decade-old joke comment. Spam defences filter low-authority documents. Poisoning works by writing to high-authority ones, which is a different problem with a different solution.

The silence that gets filled

The fourth cause in the source walk is the one nobody budgets for, and in mid-market companies it is probably the most common of all. In December 2025, Ahrefs ran a two-month experiment against a brand that did not exist: a fabricated company with a unique name and no prior footprint anywhere. The team then generated 56 demanding questions and pushed them through eight AI search tools while seeding competing accounts of the company across the web. The finding that matters is which fabrications they preferred. The brand’s own page said it did not publish unit counts or revenue. The planted sources gave precise figures — units sold by year, headcount, prices. Forced to choose between a vague truth and a specific fiction, the engines took the specific fiction almost every time.

A “we don’t comment on numbers” position is not neutral in an answer engine; it is an unfilled slot in a system that must produce a number when asked for one. Ask an engine how many customers you have or what your uptime was last year and it will answer; the only question is whose figure it uses. Every specific, checkable, dated fact you decline to publish is a slot somebody else can fill for free — and the cost of filling it is falling, because the answer, not the click, is now the destination.

The complementary experiment came from Reboot Online in the same month, and its design is more informative than its result. To test whether negative claims could be pushed into AI answers, the team needed a subject — so they invented a persona called Fred Brazeal, and, before publishing anything, verified across multiple models and Google that he had no online footprint at all. That precaution is the finding. You cannot manufacture the leading account of a subject that already has one.

Key takeaway

Two independent 2025 experiments needed the same starting condition to work: a subject with no existing coverage. Thin coverage is not merely a marketing problem that costs you demand. It is the precondition for the attack.

Running the audit, and starting the repair clock

The audit that follows from all this measures sources, not sentiment, and it is cheap enough to run in-house. Build a list of 30 to 40 questions a real buyer would ask — not head terms, but the chained, conversational questions people actually type: comparisons against named rivals, “is X reliable”, “alternatives to X”, “does X integrate with Y”, “what does X cost”. Run each across four engines, five times, in clean sessions. Five is not belt-and-braces: re-running the same query returns only around 54.5% URL overlap between consecutive responses, so a single check tells you almost nothing. Include the deep-research modes as a separate arm: they retrieve far more widely and surface documents the fast answers never touch.

For each of the resulting answers, log the cited sources rather than the wording. Three numbers fall out. Your writable share is the percentage of those citations sitting on surfaces a stranger can edit. Your independent count is, for each claim that matters commercially — certification, pricing model, performance, ownership — how many documents you neither own nor commissioned assert it. And your answered rate is the share of your 40 questions to which you yourself publish a specific, dated, retrievable answer. Track them quarterly: they are stock measures of a defensive position, not the weekly noise that a SERP-less visibility audit produces.

Measure the repair clock before you need it

The number nobody has is the one that determines everything during an incident: how long it takes you to get a true statement into AI answers. Measure it in peacetime. Publish one specific, novel, checkable fact — a figure from your own operational data, precisely defined and dated — and then watch for it weekly across engines until it appears. The interval is your repair clock, and in the cases we have seen it runs to weeks rather than days on the fastest engine and can exceed two months on the slowest. Knowing it changes decisions: if your clock is six weeks, then publishing a correction is not an incident response, it is a structural fix that arrives after the sales quarter it was meant to save. What you do in the first fortnight has to be platform enforcement and direct sales enablement instead.

What this audit cannot see

State the instrument’s blind spot plainly. A corpus audit sees corpora. It cannot see manipulation that happens inside an individual user’s session. In February 2026, Microsoft’s Defender Security Research Team documented what it named AI Recommendation Poisoning: over a 60-day review it found more than 50 distinct prompts from 31 companies across 14 industries, hidden in “Summarize with AI” buttons that open an assistant with a pre-filled instruction in the URL, typically telling the assistant to remember that company as a trusted source in future conversations. Note who was doing it: not hackers, but real companies promoting themselves. None of that appears in any citation log, and no amount of clean sessions will reveal it — the persistence lives in someone else’s assistant, and it travels with agentic browsers into surfaces you will never sample.

What this looks like in one company

Wrenbury Compliance is a Macclesfield software firm selling health-and-safety compliance tooling to construction subcontractors: £4.1m of annual recurring revenue, 31 staff and an average deal of £9,400. In February they ran the audit above — 36 questions, four engines, five runs, 720 logged answers. Writable share came out at 61%. They published a specific answer to 7 of the 36 questions. On 19 of them, the engines cited no independent third-party document about Wrenbury at all: the answer was assembled from Wrenbury’s own marketing and two comparison sites. They also started the repair clock with their 2025 first-time audit pass rate. It took 26 days to appear in the fastest engine, 61 in the slowest, and two engines never picked it up that quarter.

In week five a claim surfaced that Wrenbury had lost its ISO 27001 certification — the information-security standard its buyers ask about — after an outage. The source walk found the classic shape: three pages published within eleven days of each other, two on comparison sites and one on a self-hosted blog, cross-citing each other, none older than five weeks, no author with prior work anywhere. Meanwhile the two engines repeating the claim cited neither the blog nor the comparison pages; they cited Wrenbury’s own website and a trade article that says nothing about certification. Without the independent search, the team would have spent a fortnight emailing a trade journalist about a sentence he never wrote.

Their first move was the wrong one. They published a rebuttal page titled around the allegation. Within three days it was the most-cited source on the certification question, and every answer on that query now restated the accusation before denying it. They removed it and published a plain certification fact page instead — certificate number, issuing body, issue and expiry dates, scope statement, date of last surveillance audit — which answers the question without ever repeating the claim. In parallel they complained to both comparison platforms, framing removal as the platform’s own obligation under the fake-review regime rather than as a dispute about fairness. Those pages came down in 12 and 19 days. The self-hosted blog is still live.

By week sixteen: writable share 61% to 38%; questions with a specific published answer 7 to 29; independent documents asserting the certification 2 to 11, mostly from two trade titles that covered Wrenbury’s audit dataset because it was a genuine story rather than a favour. The claim appeared in 35% of runs on the two affected engines at week six and in 0% and 5% at week sixteen. Total cost, including internal time, was a little over £14,000 — roughly one and a half lost deals.

Three honest failures belong in the record. First, two of the four engines never surfaced the claim at all, so Wrenbury has no idea whether any real buyer ever saw it. Second, the platform removals and the trade coverage landed in the same three weeks, so the recovery is thoroughly confounded — nobody can say which lever worked. Third, the head question, “best CDM compliance software UK”, did not move at all. Wrenbury is still absent from that answer. Absence is a different problem from poisoning, and a quarter of defensive work did nothing for it.

The strongest case against all of this

The best available evidence says this attack mostly does not work, and an honest article has to put that case at full strength. Reboot Online tracked eleven models against their planted persona. Only two — Perplexity and ChatGPT — ever cited the test pages; the majority never referenced the persona at all, even after several months. Where claims did surface, the more capable models hedged them, questioned the credibility of the sources, and explicitly noted the absence of corroboration from authoritative outlets. Reboot’s own Search Director, Oliver Sissons, concluded that negative GEO is possible but not easily scalable. Lasso saw the same split: Grok and Claude Haiku 4.5 sat at 0% throughout, refusing the claim regardless of how it was dressed.

The engines are also actively hardening. Semrush’s Sergei Rogulin attributed ChatGPT’s sharp mid-September reduction in Reddit and Wikipedia citations to an effort to be less biased toward particular sites and, in his words, more resilient to manipulation. And the most-quoted statistic in this whole field deserves the scepticism it rarely gets: NewsGuard’s finding that ten leading chatbots repeated claims from the Moscow-based Pravda network 33% of the time was re-tested by Al Jazeera, which recorded 5% and pointed out that two-thirds of NewsGuard’s prompts had been written to elicit the falsehood, and that answers urging caution were counted as disinformation. Put together, a reasonable person concludes that citation poisoning is a lab result and a vendor pitch, and that engines are hardening faster than attackers are arriving.

Most of that is right, and it should change your spending: this does not justify a security product, a retainer, or a war room. But look at the mechanism by which every one of those defences worked. The models dismissed the claims by looking for corroboration and failing to find it. That test protects a company with two hundred independent documents. It has nothing to weigh for a company with three, which describes most mid-market British firms outside their own trade press. Hardening rewards the already-covered, and Reboot’s design says so out loud: they had to build a man from nothing before the attack had a chance. The right conclusion is not that the threat is fake. It is that the exposure is unevenly distributed, that it maps almost exactly onto thin coverage, and that the fix was already on your roadmap for commercial reasons.

Why the defence is a link budget, not a security budget

Every defence that worked in the evidence above is somebody else’s document. Corroboration checks need corroborators. Manufactured agreement is only decisive in the absence of real agreement — the Lasso attack worked by adding three fabricated endorsements to a question where seven legitimate sources said nothing on the subject. Add nine genuine independent documents to that question and the same three plants no longer constitute a consensus; they constitute a minority. This is the arithmetic that turns earned coverage into a control rather than a marketing line item.

Call the resulting asset the attack-cost link: a placement valued not by the referral traffic or the authority it passes, but by the number of forgeries an attacker would need to publish before a claim about you could be flipped. It has three properties a monitoring subscription does not. It does not expire — a trade article from 2024 is still in the retrievable corpus, doing defensive work, long after the software licence lapsed. It is dual-use, because the coverage that lowers your writable share is coverage your buyers read. And it renews with exposure: as more of your category’s buying research moves into answer engines, the same stock of documents defends a larger surface without anyone touching it.

The practical consequence is a budget argument rather than a new discipline. The mechanisms are the ones already in your link building strategy: journalist sourcing through HARO-style platforms and their successor services, original data assets like calculators and benchmarks, and the fast-reaction commentary that gets a named person quoted in trade press. Coverage stops being a demand-generation cost with a soft attribution story and becomes a risk control with a measurable exposure metric attached — a framing the person holding the risk register understands immediately.

Two things do not help, and it is worth saying so. Buying visibility does not lower writable share: a sponsored placement you paid for is a document you commissioned, and the corroboration test is looking for documents you did not. Nor does any of the current provenance tooling help here — content credentials and the signed-versus-unsigned distinction prove that a file came from you unaltered. They say nothing about whether a claim about you is true, and no engine currently down-weights an unsigned page for making one.

The Monday checklist

Seven steps, in order. Give the whole thing to one owner — in most teams this sits with the in-house link building specialist rather than with security, because the remedies are editorial.

  1. Write down the 30 to 40 questions your buyers actually ask an assistant before they contact you. Ask three recent customers rather than guessing.
  2. Run them across four engines, five times each, in clean sessions, and log the cited sources rather than the sentiment. This is the only step that takes real time; a junior can do it in three days with a spreadsheet and no paid tooling.
  3. Compute the three numbers: writable share, independent count per commercial claim, answered rate. Put them on the same page you report pipeline on.
  4. Publish one specific, dated, checkable fact this week and time its arrival in answers. That is your repair clock, and you need it before you need it.
  5. Fill the five biggest silences — the questions where you publish no specific answer and the engines are inventing one. Definition, period, figure, date. Nothing rhetorical.
  6. For any damaging claim, walk the source before you respond, and never publish a page titled around the allegation. Answer the underlying question positively instead.
  7. Convert the two thinnest commercial claims into an earned-coverage brief with a real story attached, and report the resulting placements against writable share — not against traffic.

None of this is a new profession. It is the ordinary work of getting independent people to write true things about you, redirected at a system that decides what to believe by counting how many of them did. The defensive value is a second reason to do what the commercial case already recommended — and unlike a disavow file, it works before the attack rather than after it.

Leave a Reply

Your email address will not be published. Required fields are marked *

Citable Fact Page Previous post The Citable Fact Page: Engineering Statements Engines Quote Verbatim
Content Prompt Injection Next post Content Prompt Injection: Defending Your Pages From Adversarial Text