TL;DR: A deep research run is not a bigger answer. It is an autonomous analyst that reads 50 to 200 sources over five to thirty minutes and writes a persistent report that circulates inside a buying committee you will never meet. The citation you earn there is a different object from an answer-box citation: its value is not the link but whether your claims survive synthesis into the body text. This article maps the five-stage report supply chain, introduces the Carry Ladder for grading your presence in these reports, and shows why third-party corroboration — not more content — is what lets your own numbers survive an analyst that cross-checks everything it reads.
The claim: more sources means the door is finally open
The optimisation advice attached to deep research modes is already settled, and it goes like this. OpenAI Deep Research fetches 50 to 200 sources per run, Gemini fetches 30 to 150, Claude fetches 20 to 100 (Presenc AI, May 2026). A standard AI answer cites three or four documents; a deep research report cites thirty or more. Therefore the retrieval door is finally wide enough for everyone: publish comprehensive, well-structured content, cover the long tail of sub-questions, and the sweep will pick you up. Your visibility tracker will confirm it, because every platform that measures AI citations folds deep research citations into the same citation-share number as everything else.
Concede what is true in this. Deep research is the widest retrieval regime that exists. Semrush measured ChatGPT running 1,130 web searches across 100 prompts under high reasoning against 245 under minimal reasoning — nearly five times the retrieval for the same questions (Semrush, June 2026) — and a full deep research run goes further still. If your page is ever going to be fetched, this is the surface where it happens.
But the settled advice rests on an accounting error. It treats every citation as the same event: an answer shown to a person at the moment of interest. A deep research citation is not shown to a person at the moment of interest. It is an input to a document — a 20-to-30-page report that gets written once, saved, forwarded, pasted into a business case, and read by people who never saw your page and never will. Counting it alongside answer-box citations is counting a footnote in a board paper as if it were a click. The rest of this article is about what the footnote is actually worth, and how to become something better than a footnote. The fundamentals have not moved — what a backlink-shaped endorsement is for has always been to let someone else carry your claim — but the carrier has changed species.
What a deep research run actually does
What is a deep research mode?
A deep research mode is an agentic AI feature that plans a multi-step investigation, runs dozens to hundreds of web searches, reads the results, identifies gaps, and returns a single long-form report with citations — taking minutes rather than seconds. OpenAI launched the first mainstream version in February 2025; Gemini, Claude, Perplexity and Grok now ship equivalents, and Google’s Deep Search produces what Google itself calls a fully-cited report.
The pipeline matters more than the branding. Every implementation runs the same four moves: it decomposes your prompt into sub-questions, sweeps the web for each one, reads and cross-checks what it fetched, and then writes. Tensoria’s May 2026 comparison describes the output as reading like a junior analyst’s briefing document, and the runtimes support the analogy — OpenAI 10 to 30 minutes, Claude 5 to 20, Perplexity 3 to 10 (Presenc AI). The retrieval budget is ten to fifty times a standard answer, which changes who gets read: the marginal source admitted to a deep research sweep is far down the tail, past the head terms, past the featured-snippet holders, into specialist pages that never win a quick answer.
Two properties of this pipeline decide everything downstream. First, the reading is adversarial: the agent compares sources against each other, and the DRACO benchmark — 100 research tasks drawn from production Perplexity usage — found citation quality and factual accuracy to be the weakest axes of every system tested, with the best achieving only 65% citation quality (Zhong et al., 2026). These systems know they are unreliable citators, and their synthesis stages are increasingly built to hedge, drop, or flag claims that only one source makes. Second, the writing is lossy: of 150 documents fetched, perhaps 30 are cited and a dozen actually shape the text. Admission to the sweep is the cheapest ticket in the building.
Google has documented its version of the pipeline plainly: AI Mode’s query fan-out breaks one question into many simultaneous searches, and Deep Search narrows the results to a small cited set in what Google calls a fully-cited report. Retrieve broadly, verify rigorously, cite sparsely — three vendors, one converging architecture. The narrowing is the point. A hundred and fifty documents enter; a dozen shape the prose; and the criterion for surviving the narrowing is not authority in the ranking sense but corroboration in the evidential sense: a moderately authoritative page whose claim is repeated across four other indexed sources beats a highly authoritative page whose claim stands alone. That inversion — corroboration outranking authority at the moment of writing — is the mechanical reason the rest of this article keeps returning to placements other people publish.
The analyst you cannot brief
Here is the axis this article runs on: deep research re-inserts an intermediating document between your content and your buyer. The field has dealt with intermediating documents before — analyst firm reports, procurement frameworks, trade-press vendor roundups — and it built a discipline for them called analyst relations: identify the analyst, brief them, correct them, feed them data under embargo. Deep research is analyst relations with the analyst removed. The analyst is instantiated fresh for every prompt, reads everything fetchable in fifteen minutes, accepts no briefings, takes no calls, and vanishes when the report is delivered. Its output is the only version of you the reader ever sees.
And the reader is worth more than the query count suggests. Forrester’s Buyers’ Journey Survey of 18,000 business buyers found 94% used AI during their most recent purchase, 55% compare vendors inside AI tools, and 47% build internal business cases with AI before any vendor contact (Forrester, January 2026). G2’s March 2026 survey of 1,076 software buyers found 41% use deep research tools regularly for software evaluations, 69% chose a different vendor than they initially planned based on AI guidance, and a third bought from a vendor they had never heard of. The reports these runs produce land inside buying committees that average 11 to 14 stakeholders on complex purchases — and 6sense’s finding that 95% of deals are won by a vendor on the day-one shortlist means the report is frequently the instrument that writes the shortlist.
So the correct mental model is not a bigger answer box. It is a bespoke, unbriefable analyst report with a print run of one buying committee — persistent where an answer is ephemeral, quoted where an answer is skimmed, and attached to a purchase decision where an answer is attached to a curiosity. Presenc AI made the compounding point directly: the reports themselves become artefacts shared inside organisations. Your citation does not expire when the chat closes. It sits in a PDF on a shared drive, doing work or damage for months.
The compression makes each report count double. Apollo’s 2026 data puts the average vendor shortlist at roughly 2.5 names, down from 3.2 a few years earlier — fewer slots, each carrying more weight, and increasingly written before any human at the vendor knows an evaluation exists. Deep research outputs are now standard raw material for the documents that formalise that shortlist: RFP question sets are drafted from them, requirements matrices inherit their comparison axes, and the business case that goes to the budget holder quotes them with the citations stripped. By the time your sales team meets the buyer, the report has been circulating for weeks, and the meeting is largely a verification exercise against what it said.
The Report Supply Chain
To act on any of this you need to know where, in the production of a report, your presence is decided. There are five stages, and they have different owners, different survival rules, and different levers. Most optimisation effort is spent on stage three, which is the cheapest stage to clear and the least sufficient.
| Stage | What happens | What decides survival | Your lever |
| 1. Brief | The user writes the prompt: the vendor category, the constraints, sometimes named candidates. | Whether your brand or category framing is already in the prompt — branded demand formed before the run. | Brand and category marketing. If the prompt names you, every later stage starts warm. |
| 2. Plan | The agent decomposes the brief into sub-questions: pricing, implementation, compliance, alternatives. | Whether a sub-question exists that your evidence answers. Several are answerable only by third parties by construction. | Publish against the sub-questions, and earn placements against the ones your own domain cannot answer. |
| 3. Sweep | Dozens to hundreds of fetches. Admission, not selection. | Retrievability and legibility: crawler access, machine-readable pages, and presence on any allowlist the run is restricted to. | Technical hygiene plus membership of the estates a restricted run trusts: trade bodies, registers, standards sites. |
| 4. Synthesis | Reading, cross-checking, contradiction handling, writing. Roughly a dozen sources shape the text. | Extractable evidence density, and corroboration: a claim made only on your own domain is hedged, flagged, or dropped. | Data-dense pages, and third-party placements that repeat your headline numbers so they survive cross-checking. |
| 5. Circulation | The report is saved, forwarded, quoted into business cases and RFP documents. | Whether your material was carried as body text. Bibliographies are not read; body text is. | None directly — circulation is downstream of absorption. Win stage four. |
Read the last column top to bottom and notice what it is: a link builder’s job description with two new line items. Stages one and two are shaped by whether independent sources already talk about you — measured entity authority formed long before the prompt was typed. Stage three adds a gate the field has never screened for, covered below. Stage four is where the strategies that earn third-party evidence stop being a ranking tactic and become the difference between being read and being carried. Stage five you cannot touch at all — which is precisely why stage four is the whole game. The same logic already governs the factors AI systems weigh when recommending products: what a machine writes down is decided upstream of the moment it writes.
Key takeaway: A deep research citation is manufactured in five stages — brief, plan, sweep, synthesis, circulation — and only the sweep is about your website. Everything decisive happens in stages you can reach only through evidence other people publish about you.
Admission is not absorption
What is citation absorption?
Absorption is when a report carries your material in its body text — your number, your definition, your comparison — rather than merely listing your URL among its sources. Machine Relations’ May 2026 analysis of 252,000 controlled trials and more than 21,000 citations found that citation breadth and citation depth diverge: the pages whose content gets deeply absorbed are a different, smaller set than the pages that get cited, and they share measurable properties — longer, internally structured, semantically aligned, and rich in extractable evidence such as definitions, numerical facts, comparisons, and procedural steps.
This divergence is the single most useful fact in the whole topic, because it means the metric everyone tracks — did we appear — is measuring the wrong stage. In a three-citation answer box, appearing is most of the value: the reader sees three links and yours is one of them. In a thirty-citation report, appearing is close to worthless: nobody reads a bibliography, and the per-citation attention collapses as the citation count grows. The value migrated into the body text, and body text is written per claim, not per page. The synthesis stage is not choosing your page over a rival’s page; it is choosing your number over a rival’s number, one sentence at a time.
And it is choosing adversarially. The agent has read everything, so every claim you make is read alongside every claim about the same quantity from everyone else. A figure that appears only on your own pricing page is, to a cross-checking synthesiser, an uncorroborated assertion by an interested party — exactly the material a system embarrassed by its 65% citation quality is being engineered to hedge. The same figure repeated in two trade publications and an industry journalist’s sourced piece is a corroborated fact. Same number, different evidential status, different fate in the report. Self-published truth loses to independently repeated truth at stage four, every time, by construction.
Watch what synthesis does when sources disagree, because that behaviour is where reports are won. The DRBench benchmark’s FACT framework — which checks whether the content at a cited URL actually supports the claim attached to it — measured citation accuracy ranging from 78% for OpenAI Deep Research to 94% for Claude with search (Du et al., 2025). Vendors know these numbers, and the engineering response has been visible across 2026: when two fetched sources give different figures for the same quantity, current systems increasingly present a range, attribute both, or flag the dispute rather than pick a side. For you, that behaviour cuts both ways. If a rival’s inflated claim stands uncorroborated next to your verified one, the report says so — the cross-check is doing your competitive work for free. But if your own two-year-old case study contradicts your current pricing page, the report flags that too, and a committee reads internal inconsistency as risk. Deep research punishes estates that disagree with themselves harder than any surface before it, which makes claim hygiene — one number per quantity, dated, everywhere — a prerequisite, not a polish.
The Carry Ladder
Because value migrates from the citation to the carry, you need an instrument that grades what a report actually did with you. Run five procurement-style deep research prompts in your category — the briefs your buyers write, not your keywords — across the engines your market uses, monthly, and grade every appearance on this ladder:
Grade 0 — Fetched, dark. Your page appears in the run’s activity log (where visible) but nowhere in the output. The sweep admitted you; the synthesis found nothing worth taking. A legibility or evidence-density failure, not a visibility one.
Grade 1 — Footnoted. Your URL is in the citation list; no claim of yours is in the body. This is the grade most citation trackers celebrate. It is worth almost nothing: bibliographies are not read by committees.
Grade 2 — Absorbed. A number, definition, or finding of yours appears in the body text, attributed. Your material now travels with the report into the business case. This is the first grade that touches revenue.
Grade 3 — Structural. Your framework organises a section: the report compares vendors on axes you published, or its cost section is built around your bands. The reader adopts your way of seeing the market without knowing it. One structural appearance outweighs a page of footnotes.
The reading: track the distribution across grades, not the count of appearances. A month that moves two footnotes to absorbed is a good month even if total citations fall.
The ladder also explains why deep research rewards a content type the answer box punishes. Long, dense, sourced, numerical pages — the ones too heavy to win a snippet — are precisely the pages whose material survives synthesis. This is the one surface where the study you spent six weeks on beats the listicle that summarises it, provided the study’s numbers exist somewhere other than your own domain. Original research assets — and of these, a dated first-party dataset is the strongest — are Grade 3 machinery: they do not just answer a sub-question, they define its axes.
The widest door has a second lock
The optimists are right about one thing: deep research is the long tail’s best odds. Moz’s 40,000-query analysis found 88% of AI Mode citations sit outside the top-10 organic results, and a deep research sweep reaches deeper still — it is the only retrieval regime in which the 150th-best page on a topic gets read at all. Specialist pages, recruitment-and-vertical-niche estates, international and non-English sources — everything the head-term auction priced out has a genuine route into these reports.
But while the door widened, a second lock was fitted. In February 2026 OpenAI shipped the ability to connect deep research to any MCP server — Model Context Protocol, the standard plug for external data sources — and to restrict web searches to trusted sites (OpenAI, February 2026). Enterprise and procurement deployments use exactly this: runs restricted to allowlists of trade bodies, standards organisations, analyst estates, and approved publications. Inside a restricted run, no amount of content quality admits you; membership does. The sweep stops being a search and becomes a roll call.
This splits the surface in two, and your plan must serve both. The open sweep rewards depth and evidence density — publish for it. The restricted sweep rewards affiliation — sponsorships, trade-body memberships, register and citation listings, standards participation: the unglamorous placements the field deprioritised for a decade because they sent no traffic. A restricted run does not care about your domain rating; it cares whether you exist inside the fence. Auditing which fences your buyers’ runs are restricted to — ask them; procurement teams will tell you — is a research task worth more than another content sprint. And note the seniority gradient: the larger the deal, the likelier the run is restricted — which means the affiliation gate binds hardest on exactly the revenue the open-sweep optimists are counting on.
Key takeaway: Deep research is simultaneously the most open retrieval regime (150 sources read, the tail finally reachable) and the most closed (enterprise runs restricted to allowlists where membership, not merit, admits you). Publish for the open sweep; affiliate for the restricted one.
What this reprices in link building
Muck Rack’s May 2026 analysis found earned media accounts for 84% of all AI citations, and deep research sharpens that arithmetic in four specific ways.
Corroboration placements — insurance for your own numbers
The highest-value placement in a deep research world is one that repeats your headline figures with attribution. A guest post that carries your median-implementation-time number, a trade-press piece that quotes your dataset, a niche edit inserting your figure into an existing sourced roundup — these are not link plays, they are corroboration plays: they change your claims’ evidential status at stage four from asserted to verified. Commission them with the number in the brief, not just the anchor text.
Practitioner testimony — the sub-question you cannot self-answer
Every procurement brief spawns a sub-question shaped like what do people who use this actually say, and it is unanswerable from your own domain by construction: a vendor’s testimonials page is an interested party’s exhibit, and synthesis treats it that way. Review platforms, community threads, and launch-platform discussions fill that role, which is why they punch so far above their authority metrics inside reports. For example, a four-comment thread where a named practitioner describes your deployment honestly — including the rough edges — routinely gets absorbed where a polished case study gets footnoted; the imperfection is the credibility. You cannot write these. You can make them likelier: ask happy customers to write where machines read, and treat an honest mixed review as an asset rather than a fire.
Evidence density as a placement screen
A brand mention in a thin listicle and a data-bearing paragraph in a dense comparison have identical value in a link audit and wildly different value to a synthesiser. Screen prospective placements for extractable evidence — will the page carry a number, a comparison, a dated finding of yours? — before authority metrics. A mid-authority page that uniquely corroborates your claim beats a high-authority page that mentions your name, on every run.
Timeliness still compounds
Deep research runs disproportionately target current questions — vendor landscapes, regulation, pricing this year — so newsjacked placements and dated event coverage enter reports while evergreen pages fight over the definitional scraps. A dated third-party account of something you did is the one evidence class that cannot be duplicated by a rival or degraded by an adversary after the fact.
Hallucinated-URL reclamation
An operational tactic almost nobody runs: deep research agents fabricate citation URLs at a measurable rate — 3 to 13% of cited URLs across commercial systems do not exist (arXiv 2604.03173, 2026). Some of those invented URLs are on your domain: plausible-looking paths to pages you never published, now sitting in saved reports inside buying committees. Mine your 404 logs for LLM-referred and direct hits to non-existent deep paths, and 301 each recurring invention to the nearest real page. It is link reclamation for links that never existed — minutes of work, and it converts a credibility-damaging dead end in a board paper into a landed reader.
Watching a surface that leaves no trace
How do I know whether deep research reports are mentioning my brand?
You cannot observe the reports directly — they are private documents on other people’s drives — so you observe two proxies: the probe distribution (your own monthly Carry Ladder runs, which sample what reports in your category currently say) and the lagged exhaust (the brand-shaped demand that report readers produce weeks later, arriving with no referrer).
The exhaust is systematically mislabelled by every analytics stack you own. Loamly’s February 2026 analysis found 70.6% of confirmed AI traffic lands in GA4 as Direct, and Similarweb estimates roughly 56% of AI-influenced visits are misattributed to search. A committee member who read about you in a report three weeks ago does not click a citation — the PDF’s links are often dead text anyway — they type your name. So the observable signature of deep research circulation is a rise in branded search and direct visits to deep product pages, offset from your placement activity by weeks, from readers your CRM has never seen. Sales teams meet the phenomenon before analysts do: prospects arriving at first contact already holding specific figures, sometimes your own, sometimes a rival’s, occasionally figures no one ever published — a hallucinated spec quoted back to you is a deep research artefact as surely as a signed contract.
Instrument what you can. Tag the phrasings unique to your published datasets and watch for them in inbound demo requests and RFP language — a committee that writes implementation risk questions in your vocabulary read a report that carried you at Grade 3. Log every won and lost deal’s first-contact knowledge level. And accept the epistemics: this surface reports to you late, anonymised, and by rumour. The probe distribution is the only leading indicator you get, which is why the monthly re-run is on the checklist rather than the wishlist.
The strongest objection: a rounding error with a fan club
Here is the hardest honest counter, stated at full strength. Deep research runs are a sliver of AI usage. Profound’s telemetry found only around 18% of ChatGPT conversations trigger any web search at all, and full deep research runs — costly, slow, capped even on paid tiers — are a fraction of that fraction. The buying-behaviour surveys measure enthusiasm, not volume. Optimising for a surface this thin is boutique work dressed as strategy, and the hours belong on the answer box where the users actually are.
Concede the volume point entirely — and then weigh instead of counting. First, the runs concentrate exactly where the money is: 41% of software buyers using deep research regularly for evaluations (G2) and 47% of buyers building business cases with AI before contact (Forrester) means the thin surface sits disproportionately under five-, six- and seven-figure decisions. A surface’s worth is spend-weighted, not query-weighted. Second, persistence multiplies each run: one report is read by a committee of 11 to 14 and quoted for months, so impression arithmetic undercounts it by an order of magnitude. Third, the direction of travel is one-way: these modes went from a $200 tier to free tiers inside a year, and the binding constraint — inference cost — only falls. Fourth, and decisively: every action in this article also improves ordinary answer-box citation. Corroborated claims, evidence-dense pages, trade-body presence, reclaimed URLs — nothing here is stranded if deep research stalls. The objection wins the volume argument and loses the allocation one.
Worked example: Ottershaw Compliance
Ottershaw Compliance is an invented but representative case: a Leeds-based vendor of FCA-compliance software for mid-tier lenders, £7.2M ARR, 58 staff, selling into exactly the procurement processes that deep research now intermediates. In March 2026 they ran the baseline: 24 procurement-style briefs (shortlist compliance platforms for a 400-seat consumer lender; compare implementation risk; build the business case) across four engines’ research modes, one run each.
The distribution, graded on the Carry Ladder: bibliography appearances in 10 of 24 reports (42%) — respectable, and exactly what their citation tracker had been celebrating. Absorbed in 2 (8%). Structural in none. Meanwhile a rival’s total-cost-of-ownership framework organised the cost section in 14 of the 24 reports: every committee reading those PDFs was comparing vendors — including Ottershaw — on the rival’s axes. Their tracker had no column for that.
Two baseline details shaped everything after. First, the run logs showed Ottershaw was being fetched far more often than cited — dark at Grade 0 in roughly half the runs where they appeared at all — so the constraint was demonstrably synthesis, not visibility, and another content sprint aimed at the sweep would have fed a stage that was already passing them through. Second, their one absorbed claim in both reports was the same sentence: a compliance-deadline statistic from a 2025 blog post that a trade publication had happened to quote. The only figure of theirs that had ever been independently repeated was the only figure that ever survived. The baseline was not just a score; it was the mechanism, demonstrated on their own estate.
Eighteen weeks of work, allocated by stage. Stage four first: they published a dated implementation-time dataset from 62 completed deployments — median 41 days, banded by lender size, re-dated quarterly — first-party operational data they already held. Then corroboration: seven placements commissioned with the numbers in the brief — two trade titles, a risk-management journalist’s sourced piece, a comparison-site methodology page among them — so the 41-day median existed in independent print. Stage three: two trade-body listings and a standards-forum membership for the restricted sweeps. Plus the reclamation sweep: nine recurring invented URLs found in 404 logs, all 301-redirected.
July re-run, same 24 briefs. Bibliography rate: 46% — statistically flat, and the honest proof that footnotes were never the lever. Absorbed: 31%. Structural: 21% — their implementation-time bands now organised the delivery-risk section in 5 of 24 reports. Inbound demos citing our research phrasing rose from 2 to 9 a month. The negatives were real: two rivals began quoting the published 41-day median in their own sales decks against them; the dataset costs roughly 11 hours a quarter to re-date; one commissioned placement sat on a site that blocked AI crawlers entirely, making it invisible to every sweep despite its authority; and the two briefs run under a client’s restricted allowlist still excluded them — the trade-body memberships had not yet propagated to the list the client actually used.
The Monday checklist
- Write five procurement-style deep research briefs in your category — the prompts your buyers write, not your keywords — and run them across the engines your market uses.
- Grade every appearance on the Carry Ladder (dark / footnoted / absorbed / structural) and record the distribution, not the count.
- List your three headline numbers. For each, count the independent domains that currently repeat it. Any number at zero is scheduled to be hedged out of the next report written about you.
- Commission two corroboration placements this quarter with the number in the brief, screened for extractable evidence density before authority.
- Identify one first-party dataset you already hold operationally and publish it dated, banded, and quarterly-maintained — Grade 3 machinery.
- Ask two friendly customers whether their procurement AI runs are restricted to an allowlist, and which one. Join what you can join.
- Mine 404 logs for recurring hits to pages you never published; 301 every invented URL to the nearest real page.
- Re-run the five briefs monthly and manage the grade distribution the way you once managed rankings.
The analyst reading your site tonight has no phone number, no badge, and no memory of your brand. It will read two hundred pages in twenty minutes, trust nothing said only once, and write the only version of you a committee of twelve will ever see — then it will vanish, and the report will not. You cannot brief it. You can only make sure that, everywhere it looks, someone else has already said what you need carried. For the wider evidence base on how citations now form, the 2026 link building statistics hub tracks the studies; the tools guide covers the monitoring stack; and if machine-readable groundwork is the gap, start with how AI systems are trained on and select sources and how agent-driven browsers change what gets read.
