TL;DR
A follow-up turn is usually not a new search. It is a new answer written from the documents the first turn already pulled in.
About 18% of ChatGPT conversations trigger any web search at all (Profound, February 2026), and follow-ups are the turns least likely to trigger one.
So your presence at turn four is mostly decided at turn one. You are not cited in the follow-up. You are carried into it, or you are not there at all.
Semrush measured this in June 2026: a brand cited at a buyer’s problem stage survived to the selection stage in 4 of 20 journeys under high reasoning, and in none of them under minimal reasoning.
Replace your single citation-share figure with two numbers: entry rate, and carry rate given entry. They point at completely different fixes.
The comparison, verification and constraint turns cannot be answered from your own domain at all, which is where earned third-party sources do the carrying.
The playbook that treats every turn as a new search
There is a well-formed piece of advice circulating about multi-turn query optimisation, and it deserves to be stated properly before it is taken apart. It runs like this. Buyers no longer ask one question; they hold a conversation. Map that conversation. Write down the chain a real customer follows from problem to shortlist to purchase, then build content that answers turn two, turn three and turn four as deliberately as you once built content for a head term. Study how people phrase a second or third question. Mark up your answers so a model can lift them cleanly. Publish the follow-up as its own addressable asset, because when the user asks it, you want to be the source the engine reaches for.
The evidence marshalled for this is real. In January 2026 Google made Gemini 3 the default model behind AI Overviews globally and shipped the ability to ask a follow-up question directly from an overview, carrying the context straight into an AI Mode conversation. Google’s own first-year usage report, published around I/O 2026, said follow-up queries in AI Mode rose by more than 40% a month in the United States, that the average AI Mode search runs around three times longer than a traditional one, and that planning and decision queries were among the strongest growth signals in the product. Users really are having conversations. The behaviour the playbook describes is not imaginary. It is not confined to Google either, which is why the question of how overviews and answers choose their sources has quietly become a question about conversations rather than about pages.
What the advice gets right, and where it breaks
Two things in it are correct. Content that answers a later-stage question does get retrieved when a later-stage retrieval happens. And the interface genuinely encourages the chain: Google built the chips, and the chips work. If you have spent the last decade building individual pages for individual queries, the instruction to extend that habit one layer deeper feels like a natural continuation of a discipline that has served you well.
The failure is not in the playbook’s picture of the buyer. It is in its picture of the machine. It assumes that each turn in a conversation runs a fresh competition for sources, in the way each query in a search session once ran a fresh competition for rankings. That assumption is wrong often enough to invalidate the plan built on it, and the direction of the error is the expensive one: it sends budget toward the turns where the contest is already over.
What actually happens between turn one and turn two
What is a multi-turn query chain? It is a single conversation in which a user asks a question, receives an answer, and then asks two or more follow-ups that depend on what has already been said. Unlike a sequence of separate searches, the turns share one context, and in most cases they also share one set of retrieved documents.
That last clause is the whole argument, so it is worth slowing down on the mechanics. Retrieval in these systems is an event, not a default state. The model decides, per turn, whether it needs the web at all. Profound’s analysis of real user conversations, published in February 2026, found that around 18% of ChatGPT conversations trigger at least one web search, a rate that held steady across the three months studied. Their reading of the turn-level pattern was blunt: follow-up turns tend to be clarifications, deeper dives, or creative tasks that do not need fresh web data, and the query that kicks off a research journey is the one worth winning.
The re-read, not the re-search
When no new search fires, the model is not working from nothing. It is working from the documents pulled in earlier, which are sitting in the conversation alongside its own previous answers. The follow-up gets answered by re-reading that set. A source that is in the set can be cited again. A source that is outside it cannot be cited at all, no matter how precisely its content matches the question just asked. This is the part the playbook has no room for: there are turns at which nothing you publish can compete, because no competition is being held.
When a search does fire, it is not your buyer’s question
Now take the other branch. A follow-up does trigger retrieval. What reaches the index is not what the user typed. Retrieval systems handle this through conversational query rewriting, a step that turns a context-dependent follow-up into a standalone search query using the conversation so far. It has to: a message like which one is cheapest carries no topic of its own. One production analysis of ecommerce assistants, published by Alhena AI in April 2026, put the share of follow-up messages containing unresolved references that depend entirely on prior turns at over 60%.
So the rewriter fills the gap from context, and the context already contains the brands, products and framings the first answer supplied. The query that goes out to be matched has the incumbents written into it. This is not a fresh contest with a level field. It is a contest whose terms were drafted by the previous round’s winners, which is a familiar problem to anyone who has run a competitor backlink analysis and noticed how much of the visible field is downstream of who was already visible.
Why the first answer sticks
There is a body of research explaining why models hold on to early positions, and it was not written with marketing in mind. Laban and colleagues at Microsoft Research and Salesforce simulated more than 200,000 conversations and found that every leading open and closed model performs significantly worse in multi-turn settings than single-turn ones, with an average drop of 39% across six generation tasks. The decomposition matters more than the headline: aptitude fell about 16%, while unreliability rose about 112%. Their qualitative account is that models make assumptions in early turns, attempt final answers prematurely, and then over-rely on those attempts, with attention concentrated disproportionately on the first and last turns of a conversation. The paper was accepted as an oral at ICLR 2026, and a follow-on study in February 2026 argued the root cause is an intent-alignment gap rather than a capability deficit.
Read that as a user and it is a defect. Read it as a source sitting inside the opening answer and it is a moat. The same tendency that makes an assistant frustrating to steer makes the evidence it grabbed first unusually durable. IBM’s MTRAG benchmark, 110 human-written conversations averaging 7.7 turns each across four domains, reaches the engineering version of the same conclusion from the retrieval side: even strong retrieval-augmented generation systems, where a model fetches documents before answering, degrade on later turns and on questions that are not standalone.
Key takeaway. Three mechanisms decide whether you are present at turn four, and only one of them is a content problem. Whether you entered at turn one. Whether the engine re-retrieves at all, which is a product setting you do not control. And whether the machine-written rewrite of the follow-up still resolves to you, which is decided off your own domain. The playbook works on none of these.
The Session Entry Map
If presence is decided at a small number of moments rather than continuously, the useful thing to own is a list of those moments. The map below sets out every point at which the set of sources in a conversation can be formed or changed, what governs it, whether an outsider can get in, and what each door actually rewards. Print it and put it next to whatever prompt list you are currently tracking.
| Moment in the session | What happens to the source set | Can a non-incumbent get in? | What that door rewards |
| Opening turn, search fires | Open retrieval on the user’s own wording, expanded into parallel sub-queries before anything is fetched. | Yes. This is the only fully open door in the session. | Broad retrievability on the category question, most of it on pages you do not own. |
| Opening turn, no search fires | No source set is created. The answer is written from model weights, and any brands named come from training. | Only if you are already in the weights. | Years of accumulated corroboration. Nothing you can ship this quarter. |
| Follow-up, no re-retrieval | The existing documents are re-read. Nothing new is fetched, and nothing is evaluated against the new question. | No. The set is closed for the duration. | Documents already inside the set that can answer more than one question. |
| Follow-up, re-retrieval fires | The follow-up is rewritten into a standalone query using the conversation, then searched. | Yes, but on a query the model wrote, containing the incumbents. | Being findable for the question as the machine phrases it, not as the buyer did. |
| User supplies the source | A paste, a link, an upload, or a brand name the user introduces unprompted. | Yes. The user opens the door directly. | Being the name a buyer already knows to type. Branded demand, not retrieval. |
How to read it
Two rows are green because an outsider can walk through them, and both sit at the extremes: the opening question, and the moment a buyer names you without being prompted. Two rows are red because they are shut, and between them they cover the majority of turns in the majority of sessions. The amber row is the only mid-conversation opening, and it is narrower than it looks, because you have to win a query written by a model that has just finished praising somebody else. For example, a facilities-management supplier that has published an excellent page on service-level agreements will not be considered at the turn where the buyer asks about service-level agreements, if the opening turn about outsourcing options was answered without them.
Entry is not presence, and your dashboard only measures the door
The industry has one number for this, usually called citation share or visibility, and it is produced by running a list of prompts and counting how often you appear. Almost every published methodology runs those prompts in fresh chats, on purpose. A widely-read practitioner guide in Entrepreneur in August 2026 recommended exactly that discipline: take your commercially important prompts, run them manually in fresh sessions across several engines each month, because a clean single-shot run is the only way to control phrasing and compare outputs side by side. As measurement design that is defensible. As a description of what buyers do it is a fiction, and the gap between the two is not random.
Semrush closed some of that gap in June 2026 with a study that tested persistence directly. They ran buyer journeys and checked whether a brand cited at the problem stage was still cited at the selection stage of the same conversation. Under minimal reasoning, no journey showed that kind of full-funnel persistence. Under high reasoning, brand continuity held in 4 of the 20 journeys tested. The same study found only about 25% of cited sources overlap between ChatGPT’s reasoning modes, and that in 51 of 100 high-reasoning responses the same domain was used more than once within a single answer.
Sit with that for a second, because it cuts in an uncomfortable direction. Carry-forward is real, it is valuable, and it is largely governed by a setting the buyer selected without thinking about it. You cannot content-market your way into high reasoning. What you can do is make sure that in the cheaper mode, where the model drops you between stages and re-retrieves instead, the rewritten query still finds you. That is a corroboration problem, and it lives off your domain.
The Entry-and-Carry Split
Stop reporting one number. Report two, plus a diagnostic.
Entry rate. The share of sessions in which you appear in the first answer. This is what your current tool measures, mislabelled.
Carry rate. Given entry, the share of subsequent turns in which you are still cited. This is what the pipeline actually depends on.
Re-entry rate. The share of sessions where you were absent at turn one and appeared later. Expect it to be small. If it is not, an engine in your category is re-retrieving aggressively and the amber row is wider than average for you.
The arithmetic. Across sessions of m turns, blended turn-level presence is roughly entry multiplied by one plus carry times m minus one, all divided by m. Two firms with 22% entry and 75% carry, and 45% entry and 20% carry, land within a few points of each other on the blended figure and are in entirely different businesses.
The read. Low entry with high carry is a discovery problem: you deserve the sessions you get and are not getting enough of them. High entry with low carry is a span problem: you get in and then fall out when the topic moves. High on both is the target. Low on both means you are debating retrieval when the honest answer is that the category does not know who you are.
Why does my citation share look stable while pipeline does not move? Because a blended figure averages the turn you enter with the turns that decide the purchase, and those move independently. A firm can hold a steady 20% while quietly losing every comparison and pricing turn in the same sessions.
One more caution before anyone rebuilds a dashboard around this. These systems are noisy at the level of a single run. AirOps found that less than 10% of the same content is cited across five consecutive runs of the same prompt, and that around 85% of pages retrieved during browse mode never appear in the final answer. Sample sizes that felt adequate for tracking rankings are not adequate here, and a change of a few points between months is usually weather. Treat it the way you would treat a sudden drop in AI citations: confirm it survives a re-run before you act on it.
There is a second sampling trap underneath the first. Entry rate is market-specific and language-specific, not a global property of a brand. Profound’s analysis of 3.25 billion citations across seven models and fourteen countries found query language to be the dominant force reshaping citation rates, with AI Overviews and ChatGPT handling non-English prompts in materially different ways. A team running English-language sessions and drawing conclusions about its standing elsewhere is measuring one door and reporting on another, which matters as much for European link building as it does across South Asian markets, where the sources engines lean on differ again.
Span: why carry is a property of documents, not claims
If a follow-up is usually a re-read of an existing set, then what determines whether you are still there at turn three is whether the document that got you in can also answer turn three. Retrieval and carry operate on documents and passages, not on your topical coverage in aggregate. A page that answers exactly one question is admitted for that question and becomes dead weight the moment the conversation moves. A page that can serve four questions in a buyer’s chain gets re-read four times.
Call that span: the number of distinct turns in a real buyer conversation that a single document can serve. It is measurable in an afternoon. Take a genuine five-turn chain from a sales call recording, and for each of your top twenty documents mark which turns it could answer without embarrassment. Most B2B sites, built under a decade of instruction to split topics into separately targeted URLs, score a median span of one.
The consolidation trade, and its limit
The obvious move is to merge. It is the right move more often than not, but it comes with a bill and a boundary. The bill is that consolidation costs long-tail rankings and sometimes destroys pages that were earning links for reasons unrelated to their retrieval value, so it needs the same care as any internal link and equity restructure. The boundary is that span is not length. Retrieval works at passage level, and analysis of 1.2 million ChatGPT answers by ZipTie found roughly 44% of citations are pulled from the first third of a page. A sprawling document with the answers buried is worse on both counts: it fails the passage test and reads badly. Span has to be built as clearly-labelled, front-loaded sections inside one document, which is a structural discipline more than an editorial one.
The span rule. Before publishing anything, name the turns it can serve in a five-turn buyer conversation. If the honest answer is one, you are buying a ticket to a single door that may not open. If it is three or more, you are buying something that keeps working after the topic moves.
The three turns you cannot answer on your own domain
Every buyer chain contains turns that your own site is structurally disqualified from answering. Comparison: which of these is better for a firm like mine. Verification: is what they claim actually true. Constraint: what happens in my specific circumstance, at my size, under my regulator. You can publish pages addressing all three, and they will be read as what they are, which is testimony from the interested party.
The engines behave accordingly, and they behave more so as they think harder. The same Semrush work found that when reasoning is switched on, user-generated sources such as forums lose roughly half their share of citations while government, academic and official documentation gain ground. The sources that survive a considered session are institutional. Muck Rack’s May 2026 analysis put earned media at 84% of all AI citations. Profound’s read of ChatGPT’s sourcing adds the pattern that matters for planning: citations travel in packs, competitors are cited side by side rather than one winner being selected, and Wikipedia sits underneath roughly one in six cited conversations as a default knowledge layer.
What to commission, and why it carries
Third-party documents have naturally high span for a buyer chain, which is why they persist. An independent category comparison feature answers the comparison turn, the pricing-model turn and often the integration turn from one page. A sector register answers the verification turn and the constraint turn at once. Neither is achievable on your own domain at any budget. This is the unglamorous half of modern link building, and it looks nothing like the placement lists most tools generate, because the selection criterion is not authority score but how many turns of a conversation one document can hold.
In practice the shortlist is short and boring: accreditation and membership registers, trade-body directories, standards and guidance pages, independent category round-ups and comparison listicles, and the sector title that publishes an annual buyer’s feature. Add to that the first-party assets that give an independent writer something to cite: a published dataset, an openly documented data feed, or a genuine calculator with a stated method. Owned-but-off-domain surfaces such as public documentation hosted on third-party platforms sit awkwardly between the two: they carry your framing, so a model discounts them as it discounts your own site, but they are often retrievable where your site is not. Community sources such as technical forums and news aggregators still matter for the opening turn, but the Semrush finding says plainly that their share decays exactly where the money is.
Worked example: Penbury Veterinary Software
Penbury is an invented but deliberately specific case: a Leamington Spa practice-management software firm, 4.6 million pounds of annual recurring revenue, 310 UK veterinary practices, an eleven-person marketing and product-marketing team. Their buyers are practice owners and managers who research in one sitting, usually in the evening, usually in a single conversation.
What the session protocol found
Through 2025 Penbury tracked 45 prompts monthly in fresh chats and reported a citation share of 24%. In February 2026 they rebuilt the test. They took twelve recorded sales conversations, extracted the real question chains, and built 60 five-turn sessions rather than 300 independent prompts. The results split the old number in half. They entered 13 of 60 sessions at turn one, an entry rate of 22%. Given entry, they were cited in an average of 2.9 of the four following turns, a carry rate of 73%. Of the 47 sessions they missed at turn one, they appeared later in 4, a re-entry rate under 9%.
The blended turn-level figure came out at 18%, close enough to the old 24% that nobody would have investigated. The session view said something the prompt view could not: in 78% of buyer conversations Penbury was absent for the entire decision, and the pricing and shortlist turns in the sessions they did win were being answered from documents selected before price had been mentioned. Their highest-carrying document was not theirs at all. It was a veterinary trade title’s annual practice-software feature, which answered the comparison turn, the integration turn and the pricing-model turn from a single page.
What they changed
Over fourteen weeks they did three things. They consolidated 190 thin pages into 96 documents built for span, with front-loaded sections per turn. They published two first-party assets an outsider could cite: a dated integration matrix covering named laboratory, insurer and payment systems, and a costed migration-time study drawn from their own onboarding records. And they commissioned 18 placements chosen by span rather than authority score, including two buying-group directories, three trade titles, a continuing-education provider listing, two independent comparison sites and eight regional practice-network pages, paced deliberately across the quarter rather than dropped in one burst.
What it produced, and what went wrong
At sixteen weeks: entry rate 22% to 34%, carry rate 73% to 78%, average turns cited per entered session 2.9 to 3.4, demo requests 34 to 41 a month. The negatives are worth more than the positives. Consolidation cost roughly 60 long-tail rankings and organic clicks fell 8%, which had to be defended in two consecutive board meetings. Each session run costs about eleven analyst hours, so the protocol is quarterly, not monthly. A re-run four weeks later with no work in between moved entry rate by five points in both directions across engines, confirming that a single reading proves nothing. And one commissioned comparison placement was quietly updated by its publisher four months later to include a competitor, after which Penbury’s share of comparison turns in the sessions it fed fell measurably. An earned document is a live asset owned by somebody else, and it can be edited without telling you.
The strongest case against this argument
The best objection is not that carry-forward is unreal. It is that carry-forward is a temporary artefact of cheap defaults that the industry is actively engineering away. Stated at full strength: Perplexity searches on essentially every prompt, agentic browsers fetch continuously, reasoning modes issue fresh tool calls per turn, and practitioner analyses of AI Mode report that it cites ten or more sources per answer and then cites a different set again on the follow-up. On that trajectory every turn becomes an open competition, the closed re-read disappears, and the sensible plan is precisely the one dismissed above: build content for each turn in the chain and be there when the engine comes looking. It is a serious position held by serious people, and it shares its logic with the argument about what an agent-driven fetch is actually worth, which also assumes the machine looks again every time.
That objection is directionally right about where the products are heading, and it should be conceded rather than minimised. Four things bound it.
- It kills the playbook faster than it kills this argument. If every turn re-retrieves independently, your turn-three page does not get carried in either. It has to win a fresh retrieval on a machine-written query, against the whole web, on a question where a comparison or verification claim from your own domain is discounted. Full re-retrieval makes corroboration more decisive, not less.
- Fresh retrieval is not fresh competition. The rewrite is generated from a context containing the incumbents’ names. The set changes; the framing does not reset.
- The current numbers describe a mixed regime, not the future one. Around 18% of ChatGPT conversations trigger any search at all, and persistence held in 4 of 20 journeys under high reasoning and none under minimal. Both mechanisms are live simultaneously, and no engine is in the state the playbook assumes.
- The recommended actions are invariant to the outcome. Enter at the opening question, and be corroborated off-domain so a rewritten query still resolves to you. Those hold whether re-retrieval goes to zero or to every turn, which is the honest test of a strategy rather than a forecast.
Three findings would damage this argument, and they are worth writing down before they arrive. First, panel data showing that membership of the cited set at turn four is statistically independent of membership at turn one. Second, a controlled test in which a page absent at turn one, but matched precisely to the turn-four question, is cited at turn four as often as it is when that same question is asked as a fresh standalone prompt. If those rates are equal, the door is always open and this whole framing is wrong. Third, an engine shipping per-turn retrieval as an audited default. The second is runnable by any reader with a fortnight and a spreadsheet, and it is the test this argument should be judged on.
What to do on Monday
Nothing here requires a new platform, and most of it is a rearrangement of work you are already doing. The sequence matters more than the speed.
- 1. Pull five real buyer conversations from recorded sales calls and write out the actual question chains, in the buyer’s words, in order. Do not invent them from keyword tools; the whole point is that the wording is no longer observable from the outside.
- 2. Convert twenty of your tracked prompts into twenty five-turn sessions and run them once. Record, per session, whether you entered at turn one and at which later turns you were still cited.
- 3. Report entry rate and carry rate as two separate lines from now on. Retire the blended figure, or keep it clearly labelled as an average of things that move independently.
- 4. Score your top twenty documents for span against those chains. Any document scoring one is a candidate for merging into a document that scores three.
- 5. For every turn you score as unanswerable from your own domain, write down which specific third-party document would answer it. That list, not a generic prospect export, is your outreach plan for the quarter.
- 6. Commission for span. Prefer one register or comparison feature that answers three turns to three placements that answer one each, even when the second option has better headline metrics.
- 7. Re-run the sessions once before you conclude anything, and check both the engine and the reasoning mode you tested. Two of the numbers in this article change entirely with that setting.
- 8. Diarise a quarterly re-read of your commissioned placements. They are documents on somebody else’s site, and they change. Whoever owns your AI visibility reporting should own that check too, alongside whoever owns outreach and relationships.
The instruction to optimise for the follow-up is not wrong because follow-ups do not matter. They matter enormously; they are where the money is. It is wrong because it describes an act with, in most sessions, no moment at which it can be performed. The commercial turns in a buyer’s conversation are answered from a set of sources chosen at the moment of least commercial intent, by a machine that is unusually reluctant to change its mind. You do not win the follow-up. You arrive before it, with documents that keep answering after the subject changes, and with third parties saying the things you are not permitted to say about yourself.
