TL;DR
The claim: AI systems lift whole paragraphs from pages and drop the attribution that came with them. The quotable-soundbite advice of the last two years optimises for being extracted, which is the easy half. The hard half is surviving compression with your name still attached.
The rule: if deleting your name leaves the sentence true, your name will be deleted. Attribution stated as a frame — “according to us”, a byline, a source line under a chart — is metadata, and metadata is the first thing a machine discards.
The method: run the compression test on every fact you want quoted, classify the result as loose, hollow, welded or inert, then move the identifying token from the frame into the assertion using a named measure, a coined term or a bound qualifier.
The payoff: a welded sentence produces the inherited citation — your name travelling inside copies you never made, through third-party restatements, and returning as branded search and unsolicited links.
The soundbite advice is optimising the wrong property
The dominant content advice for AI search in 2026 runs roughly like this: add statistics, add quotable one-liners, write a crisp 30-word answer under every heading, and the machines will quote you. It is repeated with unusual confidence, and it does have a real academic origin. The Princeton GEO paper — GEO being generative engine optimisation, the practice of editing pages to change what AI answers say — tested nine content modifications across a 10,000-query benchmark and found three that worked: adding statistics, adding credible quotations and citing sources each raised the study’s position-adjusted word count metric by roughly 30 to 40 per cent. Keyword stuffing, tested alongside them, came out around 8 per cent below the unmodified baseline.
That paper (Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, KDD 2024) is good work. The problem is what happened to it afterwards. Two later benchmarks tested the same family of tactics under harder conditions and found much less. C-SEO Bench, from Puerto, Gubri, Green, Oh and Yun at NeurIPS 2025, evaluated nine conversational-search optimisation methods across six domains and two tasks, and added the variable the original study left out: what happens when more than one page in the candidate set adopts the tactic. Their finding was blunt. Most of the methods were largely ineffective, several actively reduced a document’s standing, and simply moving a source higher in the retrieved context was around 7.6 times more effective in their retail measurement than the best optimisation method they tested. They also observed a congestion effect — as adoption rises, the gains dilute toward zero.
A 2026 follow-up went further. The FeatGEO study tested the token-level heuristics on three engines and found them falling below the unmodified page on all three: from 13.34 per cent to as low as 10.92 on GPT-4o-mini, from 8.89 to 4.62 on Gemini, and from 5.20 to 2.75 on Qwen-plus. The authors’ reading is worth sitting with: isolated text-level edits are not reliably useful and can disrupt the natural writing patterns that models prefer to cite in the first place. Sprinkling stat-shaped sentences through a page is now the tactic with the weakest evidence behind it, not the strongest.
There is a simpler reason the advice underperforms, and it has nothing to do with benchmarks. A snappy, elegant, 30-word claim is by construction the easiest thing in the world to say in twelve words instead — and when a machine says it in twelve words, the words it drops are the ones identifying who established it. The property that matters is not quotability but what survives when someone else compresses your sentence. Those are different targets, and most pages aim at the first.
What is a citable fact page? It is a page built around a small number of statements engineered so that a machine or a human restating them cannot easily remove the identifying token — the named measure, coined term or bound condition that makes the claim what it is. It is not a statistics page or an FAQ. Those are formats. A citable fact page is defined by a property of its sentences, not by its layout, and a page can carry the format perfectly and fail the property completely.
What a machine quotation actually is
Before engineering anything, it is worth being precise about the physical event we are trying to influence. Machine extraction is not new — featured snippets ran the same experiment on the open web for a decade — but it is now far more observable than most people realise.
Click a citation inside Google’s AI Mode and you often do not land at the top of the page. The URL carries a text-fragment directive — the #:~:text= syntax, a browser feature Chrome shipped in 2020 that scrolls to and highlights a specific run of text without any cooperation from the page author. It is the same plumbing behind the citation chips in AI Overviews. The engine is not pointing at your document. It is pointing at a passage inside your document, and telling the browser exactly which one. That highlight is a receipt. It tells you, without ambiguity, which words were taken.
A year-long study by the content platform Pillarbase, published on Search Engine Land in July 2026 as sponsored content, counted 15,699,298 AI Mode citations across 148 industries and found that 47.7 per cent of them were scroll-to-text highlights rather than plain links. Four findings from that dataset matter here, and the vendor provenance is worth keeping in view — the study was written to sell a product, and Search Engine Land neither confirms nor disputes its conclusions.
- Engines lift paragraphs, not one-liners. The median highlighted passage was 117 words — a complete, multi-sentence answer, not a soundbite.
- Position inside the paragraph decides everything. Around 80 per cent of extracted passages put the answer in the first sentence.
- Self-containment is close to mandatory. About 85 per cent of highlighted passages made sense without the surrounding page.
- Reuse is extremely concentrated. 80.9 per cent of passages were cited once, but roughly 2,300 were reused 61 times or more, and the single most-recycled passage was cited 661 times across 483 distinct queries.
The same dataset reports one finding that sits awkwardly against the prevailing advice: standalone statistics were rarely reused across queries. The passages that got recycled were question-led explanations, not stat blocks, which is consistent with what studies of AI Mode source selection keep finding.
There is also a hard link back to ordinary ranking work. In that dataset, pages with one to four highlighted passages had a median organic position of 11; pages with 21 or more had a median position of 1, with 67 per cent ranking first outright. Among passages reused a hundred times or more, 76 per cent came from pages ranking first. This is correlation and the author says so plainly, but it means the sentence-level work described here sits on top of, not instead of, the link building strategies that get a page into the candidate set at all. If nothing retrieves you, nothing quotes you.
Key takeaway
Extraction is passage-level and largely solved: write a self-contained 75–150 word answer, lead with the answer, put a literal question in the heading. That gets you lifted. It does nothing whatsoever to ensure that the lift carries your name — and in AI Mode, the friendliest surface there is, the text fragment at least links back. In a chat answer, a summary, a repost or a human writer’s rewrite, there is no fragment and no link.
The compression test
Machine restatement is lossy, and it is lossy in a predictable direction. It keeps the proposition and discards the frame.
The best evidence for this comes from a large 2026 audit of AI Overviews by Haofei Xu, Umar Iqbal and Jacob Montgomery at Washington University in St. Louis, who issued 55,393 trending queries over forty days, decomposed the answers into 98,020 atomic claims and checked each against the pages the answer itself cited. A claim was scored clear when the source supported it literally — 84.61 per cent of claims. It was scored vague when the source supported it only in broader or less specific terms — 4.36 per cent. The remaining 11.03 per cent were unsupported, and omission dominated: a claim no cited page mentioned at all was about 2.6 times more common than a claim a cited page contradicted.
That vague category is compression damage with a name. One claim in roughly twenty-three arrives at the reader having lost the specificity that made it worth publishing — which is why the link building statistics most often repeated in the wild are the ones stated in the fewest words. Ahrefs found the same instability from the other direction in 2026: re-run the same AI Overview query and only about 54.5 per cent of cited URLs and 54 per cent of named entities persist between consecutive responses. The wording churns constantly while the meaning stays put. Entities and propositions survive; phrasing does not, which is one reason entity authority behaves so differently from page authority in these systems.
So here is the test: take any sentence you want quoted and rewrite it as short as you can without losing information. Whatever survives is the part a machine will emit; whatever you were able to cut is the part a machine will cut too. It takes about ninety seconds per sentence and sorts every claim you publish into four classes.
| Class | What survives compression | What happens to your name | What to change |
| Loose | Almost nothing. The whole sentence restates in half the words with no loss. | Deleted. There was nothing holding it in place. | The claim is generic. Either add something only you can supply, or stop competing for it. |
| Hollow | A number or a date. The frame around it goes. | Deleted. The figure travels naked, as “studies show” or “around 60%”. | Bind the number to a condition or definition that has to travel with it. |
| Welded | A named measure, term or condition that the claim cannot be stated without. | Carried. Removing it makes the sentence vaguer or false. | Nothing. This is the target state. |
| Inert | Nothing gets lifted at all — the sentence is hedged, conditional or spread across three clauses. | Irrelevant. You were never quoted. | Split it. One proposition per sentence, answer first. |
Work one fact through all four classes. An independent testing laboratory has established that a class of sealed glazing unit loses seal integrity faster than the industry assumes.
- Loose: “Our research shows that sealed units can degrade faster than expected in real conditions.” Compresses to “sealed units degrade faster than expected”. Nothing identifies anyone.
- Hollow: “According to our 2026 study, 6 per cent of units failed early.” Compresses to “about 6% of units fail early”. The number lives; the study, the year and the laboratory do not.
- Welded: “Under the Haverstock 90-minute soak test, 6.2 per cent of laminated units failed the seal-integrity threshold (n = 312, 2024–2026).” You cannot compress out the test name without making the figure meaningless, because the figure is defined by the protocol.
- Inert: “While results vary by installation context, our data would suggest that a minority of units may, in some circumstances, fail earlier than projected.” There is no span here worth lifting.
The hollow and welded versions are both true and both contain a number. Only one makes the identifying token load-bearing.
The deletion rule and the three welds
If deleting your name leaves the sentence true, your name will be deleted. That is the whole rule, and it explains why so much careful attribution work produces nothing — including work done by people who understand how link building actually works very well. Most attribution on the web is positioned as a frame around the claim rather than as part of it: an “according to” clause, a byline, a source line beneath a chart, a citation in a caption. All of those are metadata — statements about where a fact came from rather than statements of the fact. Compression removes metadata first, because dropping it costs the summariser nothing in accuracy. The engine is not being unfair. It is doing exactly what a good sub-editor does.
There are three ways to move the identifying token out of the frame and into the assertion.
1. The named measure
You define the instrument, and the instrument carries your name. A soak test, an index, a rating scale, a severity band, a diagnostic. Once a figure only means something relative to a defined protocol, the protocol name is not decoration — it is the unit. Strip it and the number becomes uninterpretable, so the compressor keeps it. It is the strongest weld and the most demanding: the measure must be real, documented and repeatable. A protocol nobody can follow is a marketing device, and it will be treated as one by exactly the people whose links you want. The discipline is closer to publishing content credentials than to copywriting.
2. The coined term
You name the phenomenon rather than the measurement. If a category genuinely lacks a word for something and you supply one that earns its keep, restating the fact requires the word, and the word carries provenance for as long as it stays associated with you. To test whether a coined term is doing work, try to state the fact without it. If the sentence becomes noticeably longer or clumsier, the term is load-bearing. If it becomes identical minus a piece of jargon, you have added a synonym, and it will compress away with everything else.
3. The bound qualifier
You state the number with the condition that makes it true, and the condition is one only your position lets you specify — a population you can define, a period you can bound, a threshold you set. “6.2 per cent” travels alone. “6.2 per cent of laminated units at the 90-minute threshold” cannot, because the threshold is doing the arithmetic. It is the weakest of the three, and most likely to erode as the qualifier becomes industry shorthand. It is also the easiest to deploy tomorrow, needing nothing more than an editing pass and whatever technical groundwork already makes the page fetchable, which is why it is usually the right place to start.
Key takeaway
None of this works if the underlying thing is not yours. Naming a measure you did not build is the one move here that actively costs you: the first specification writer who checks finds the real source, and what you have manufactured is a brand-accuracy problem propagating into answers you cannot edit. The weld only holds when there is something real behind it.
Writing the page
A citable fact page is a short document doing one job. The shape below is assembled from what the extraction evidence actually supports, not from a template.
The fact block
One heading, phrased as the literal question a person would type. Beneath it, a single paragraph of roughly 75 to 150 words that opens with the answer in the first sentence and then supports it. Inside that paragraph — not in a caption, not in a table, not in a footnote — put the four things that make the claim checkable: the definition of what was measured, the period it covers, the sample size, and the date it was last verified. This matters more than it sounds, because a text fragment stores only the opening and closing words of a contiguous run. Anything sitting outside the paragraph is outside the quotation. A source line under a chart is, for these purposes, invisible.
Then stop. The instinct to add a summary box and an FAQ restating the same figure is the most common way teams sabotage this. Each restatement is another candidate span competing for the same lift, and the version that usually wins is the shortest one — which is almost always the hero-banner version, written by someone optimising for visual impact, with the qualifier and the method name stripped out for the sake of the line length. If you are going to state the fact more than once on your own site, every instance must carry the weld. The alternative is to state it once.
What to do with the rest of the page
Everything else on a fact page is support: how the measurement was taken, what it excludes, what changed since the last edition, who to contact to query it. It rarely gets quoted and should not be written as though it will be. Its job is to make the quoted sentence survive scrutiny when a journalist or standards committee checks — which is the moment the link gets placed. A fact that cannot survive a phone call earns one citation and no second one.
Measuring it: the quoted-span audit
Nothing above is worth doing on faith, and the measurement is cheap. The audit answers two questions that are usually collapsed into one: are we being lifted, and do the lifts carry our name?
Assemble twenty to forty prompts that a real buyer would ask and that your fact pages genuinely answer — the same sampling discipline you would apply to competitor backlink analysis. Run them across the engines that matter, capture the answer text, and compute two things: the longest run of consecutive words shared with your page, and whether your identifying token appears inside that run. Then report a lift rate — the share of answers containing a run of five or more consecutive words from your page — and a welded rate — the share of those lifts that carry the name. Five words is not arbitrary: it is the threshold SE Ranking used when it found that 69 per cent of the AI Overviews in its media-presence research reproduced a fragment of five or more consecutive words from a source.
def longest_span(answer, source):
a, src, best = norm(answer).split(), norm(source), 0
for i in range(len(a)):
for j in range(i + best + 1, len(a) + 1):
if ” “.join(a[i:j]) in src: best = j – i
else: break
return best # words; norm() lowercases and folds quotes, dashes, spaces
Interpret the number against a floor rather than against zero. When the authors of QUIP-Score — a metric built to measure exactly this kind of overlap between generated text and a source corpus — scored random text that was not quoting anything, it still returned around 17 per cent n-gram overlap, against 99.9 per cent for documents that were verbatim copies. Short matches happen by accident. Treat runs of four words or fewer as noise, five to twelve words as a lifted core, and thirteen or more as the sentence having travelled intact.
Cost, failure modes and what would make me abandon it
At forty prompts across four engines run weekly, this is 640 captured answers a month. The model calls are negligible; the real cost is capture engineering, a few days of work once, and it slots into whatever link building tools stack you already run for AI brand monitoring. Four failure modes will bite you, and all four are avoidable.
- Run-to-run volatility. The same prompt returns different wording on different days, so a single run tells you almost nothing. Repeat each prompt at least three times per capture window and report the median.
- Character normalisation. Curly apostrophes, en dashes and non-breaking spaces silently destroy exact matching. Fold them before comparing or your lift rate will read as zero for reasons that have nothing to do with your writing.
- Attribution outside the span. An engine may name you in a sentence adjacent to the lift. Score that separately — it is a real outcome, but it is not what the weld is for, and conflating the two will make a failing page look like a working one.
Record the capture date, engine and version, prompt set, locale and page revision with every run. Without that metadata a six-month comparison is meaningless: you will have changed the pages and the engines will have changed themselves.
Failure threshold. Run it for two quarters. If the lift rate is healthy but the welded rate stays near zero after a rewrite, the problem is the sentence and the fix is another pass at the weld. If the lift rate itself is near zero, the problem is not your wording at all — you are not in the retrieved set, and the honest response is to stop editing sentences and go back to earning the authority and coverage that put a page in the candidate pool. That is what the multi-actor benchmark evidence points at, and it is a cheaper diagnosis than a year of rewriting.
A worked example
Haverstock Testing is a hypothetical but deliberately ordinary business: an independent materials-testing laboratory in Sheffield, £6m turnover, 34 staff, testing sealants, coatings and glazing units for construction clients. It publishes technical notes because its engineers like writing them, and it has never treated them as a marketing asset.
Week 0: the audit
The audit found fourteen fact-shaped sentences across nine technical notes. Running 32 prompts three times each across four engines produced a lift rate of 21 per cent — respectable, and evidence the pages were being retrieved and read. The welded rate was 4 per cent. One lift in twenty-five carried the laboratory’s name inside the quoted span. The rest had been absorbed as anonymous industry knowledge: “testing has shown that roughly 6 per cent of laminated units fail early”, with no source Haverstock could point at. Every one had been written in the hollow pattern, with a perfectly good “according to Haverstock Testing” clause in front of a number that travelled without it.
Weeks 1–4: the rewrite
The laboratory already ran a proprietary accelerated-ageing protocol, known in-house as “the 90-minute soak”. It had never been named in public or described in a way an outsider could cite. Naming it, publishing the method, and rewriting nine notes so that every headline figure was stated relative to the named test took four weeks — one engineer at roughly two days a week, plus a technical editor. No new research was commissioned. The facts were unchanged. Only the sentences moved.
Week 12: results, including the ones that did not work
Lift rate rose from 21 to 34 per cent, roughly what the extraction literature predicts from restructuring paragraphs to lead with the answer. Welded rate moved from 4 per cent to 46 per cent. Eleven referring domains arrived without outreach: four specification writers, two trade publications, a university reading list, a standards-committee working paper and three supplier pages. Nine anchored on the protocol name rather than a generic phrase.
Three things did not work, and they are the useful part. The head term the commercial team cared about — a high-volume category phrase — did not move at all, because a fact page is not a category page and no amount of sentence engineering makes it one. Two of the four engines never quoted the laboratory at any point in the quarter, and the pattern of which two changed between weeks 6 and 10, which is a warning against reading engine-level differences as strategy. And the biggest single traffic movement in the period came from an unrelated technical fix to a rendering problem on the notes template, which makes the quarter’s numbers confounded and means the honest read is directional. A clean comparison needs a second quarter with nothing else changing.
The strongest objection to all of this
Here is the argument against, made as well as I can make it.
You are engineering for a behaviour that is being deliberately removed. Verbatim reproduction is precisely what publishers are suing over — Chegg and Penske Media against Google, Dow Jones against Perplexity, a complaint from 300 French newspapers over AI Overviews in August 2026 — and the obvious defensive response for any engine is to paraphrase harder and quote less. Google’s own June 2026 generative-AI guidance tells site owners they need no AI-specific rewrites at all, because these features run on core ranking. So you are proposing sentence surgery for a copying behaviour that legal pressure will engineer away, on the advice of an engine that says not to bother.
The first half of that is correct, the litigation trend is real, and it is reshaping how sources get used for training and retrieval. But it argues for the technique rather than against it, because the weld is not built for copying. It is built for paraphrase.
Paraphrase preserves propositional content and discards framing — that is the definition of the operation. A hollow sentence loses its attribution under paraphrase because the attribution was framing. A welded sentence keeps it, because the identifying token is part of what the sentence asserts. As engines quote less and restate more, the gap between those two constructions does not narrow. It widens. Verbatim quoting is the easy case, where a fragment links back whether or not you did anything clever. The paraphrased case is the one where sentence construction is the only thing carrying your name, and that is the case that is growing.
Two honest concessions go with that. First, if you have no measure, no term and no bounded population of your own — if your pages restate facts anyone could restate — there is nothing here to weld and this article is cosmetic advice. Establish something first. Second, a page whose every sentence says your own name reads as branding spam to a human, and human readers are the ones who place links. The discipline is one welded statement per page and ordinary prose everywhere else. If a section reads like a press release you have overshot, and will lose the authenticity that made the page worth citing.
The link this earns: the inherited citation
The payoff of a welded sentence is a specific and slightly unusual kind of link.
When a fact is welded, it propagates through copies you never made. A supplier writes it into a spec sheet. A forum answer restates it. A trade journalist uses it as background. A competitor’s comparison page cites it against themselves because it is the only figure available. In each the identifying token comes along, because it is part of what they are saying — not because anyone chose to credit you. That is the inherited citation: your name travelling inside other people’s sentences, at their expense rather than yours.
The links follow the name rather than the sentence, and usually with a lag. Someone encounters the protocol name in a third-party document, searches it to find out what it is, and lands on you — which converts an unlinked mention into a branded query and then into a link on their next update. This is why the distinction between a citation and a link matters operationally here: the citation arrives first, unlinked and often unnoticed, and the link is the second-order effect. Teams that measure only referring domains in month one conclude the approach failed.
What makes it worth building is the renewal property. A canonical-answer page renews as more people install or adopt whatever it documents. A statistic renews when you publish the next edition. An inherited citation renews through diffusion — every restatement by a third party creates a fresh surface carrying your name, and those are themselves restated. The asset compounds without further publishing effort, which is the opposite of most AI-visibility tactics, where the work has to be repeated to keep the result.
Set expectations accordingly. Volumes are modest — eleven domains a quarter, not a hundred. The domains skew technical and unglamorous: specification writers, trade bodies, procurement guides, university reading lists. It will not move a contested commercial head term, and it is not a substitute for earning coverage or for the ordinary work of becoming a source AI systems recommend. What it does is stop the leak — the situation, which is the normal one, where your work circulates widely and anonymously and returns nothing.
Your Monday checklist
Seven steps, in order, each finishable in a working day or less.
- List every fact-shaped sentence you publish. One spreadsheet, one row per claim, one column for the page it sits on. Most teams find between ten and forty.
- Run the compression test on each. Rewrite as short as possible without losing information. Mark each row loose, hollow, welded or inert.
- Find your measure. Ask the people who do the work what protocol, threshold or definition they use internally and have never named publicly. There is almost always one.
- Name it and publish the method. A separate page describing how the measurement is taken, what it excludes and who to contact. Without it the name is unusable by anyone careful.
- Rewrite the top five claims. One paragraph each, 75 to 150 words, answer first, literal question in the heading, definition and period and sample size inside the paragraph.
- Delete or weld every duplicate. Hunt down the hero-banner and summary-box versions of the same claim. Either bring them into line or remove them.
- Baseline the audit before you publish. Capture lift rate and welded rate on the old wording first, or you will have no way of telling whether anything worked.
The measurement discipline behind all of this is the same one that governs any serious claim about AI visibility: baseline first, repeat runs, report medians, be honest about confounds. The engineering is the easy part; the temptation to declare victory on one lucky screenshot is the hard part.
