TL;DR
— There is no single “recommendation graph”. Three different memories decide whether an agent puts your name forward, and they update at wildly different speeds and are written by wildly different people.
— The layer most visibility budgets target — retrieval consensus — has almost no memory. Two repeat ChatGPT answers to the same prompt shared only 21.2% of their cited domains across 693,509 answers (Parse, March–April 2026). You do not hold a position there. You hold a rate.
— The layer that does remember — the delegated memory inside a buyer’s own assistant — is genuinely persistent and genuinely compounding, but it is written by outcomes, not by content, and there is no way to buy into it.
— Two instruments: the Three-Memory Model, which sorts the layers by who writes them, and the Contestable-Share Calculation, which prices what a citation is actually worth once repeat demand stops re-opening the question.
— The uncomfortable division of labour: marketing wins the audition, operations wins the contract. Most 2027 plans fund only the first half.
1. The graph everyone is selling you a seat in
A specific phrase has entered pitch decks over the past year: the AI recommendation graph. It is used the way “the knowledge graph” was used in 2018 — as a structure that exists somewhere, that your brand is either inside or outside of, and that rewards early entry with a compounding position. The pitch that follows is always the same shape. Get into the graph now, while the category is unformed, and the assistant will keep returning your name because it has learned to prefer you.
The pitch is built on two observations that are both true and that pull in opposite directions.
The first is that assistants now hold durable memory about individual users. OpenAI’s June 2026 memory update, which it calls Dreaming, lets ChatGPT revise stored memories as time passes rather than accumulating a static list, and gives users an inspectable memory summary. That is a real persistence layer, and it plainly affects what gets recommended.
The second is that answers are extraordinarily unstable. Re-run the same buyer-intent prompt and the set of sources behind the answer largely changes. Parse measured 16,143 ChatGPT prompts re-run roughly 22 times each between 26 March and 25 April 2026 — 693,509 answers in total — and found that any two repeat answers to the same question shared only 21.2% of their cited domains. Google AI Overviews was steadier at 31.5%, which is still two-thirds churn.
Both facts get used to sell the same product, which should be the first warning. If the system remembers, why is it unstable? If it is unstable, what exactly are you buying a position in?
The answer is that “the recommendation graph” is not one thing. It is a word doing the work of three separate mechanisms with three separate owners, three separate update speeds, and three separate ways of being won. Collapsing them produces two specific and expensive errors, which this article is about: treating a flow as if it were a stock, and treating an operational output as if it were a marketing target.
What is an AI recommendation graph?
It is an informal term for whatever makes an assistant return the same brand names repeatedly, and it has no single technical referent. In practice it describes three distinct stores: the brand priors already baked into a model’s weights, the consensus a retrieval system assembles from the live web at query time, and the per-user memory an assistant keeps about a specific buyer. Only the third is a stored preference in any literal sense.
The distinction is not academic. It determines which team in your business owns the outcome, what a citation is worth, and which of your current metrics is measuring something that does not exist.
2. Three memories wearing one name
Sort the mechanisms by the only question that reliably separates them: who or what writes this, and how fast does it change its mind?
The parametric prior: what the model already believes
This is brand knowledge encoded in the weights during training. It is why an assistant can name five vendors in your category with no retrieval at all, and why the names it produces skew toward the large and the long-established. It updates only when the model is retrained, which means it lags the market by months and cannot be edited, appealed, or corrected on any timescale a campaign operates on. Its practical significance is narrow but real: it decides who is in the room when the system answers without looking anything up.
For example: ask an assistant with browsing disabled to name UK commercial catering equipment suppliers and you get a stable, conservative list drawn entirely from this layer. Turn browsing on and the list moves. The gap between those two answers is a rough read on how much of your presence is earned recently versus inherited from the training corpus, and it is a cheap diagnostic almost nobody runs.
Retrieval consensus: what the open web agrees today
This is the layer nearly every visibility budget targets. At query time the system assembles sources, reads what they say about the category, and composes an answer. Nothing about your brand is stored here between queries. The consensus is reassembled from scratch each time, from whatever the corroboration set currently contains. That is why the same question produces different sources an hour later, and why a screenshot proves nothing.
Delegated memory: what this buyer’s agent has learned
This is the genuinely persistent one — the per-user store inside an assistant, and increasingly inside procurement and replenishment agents operating under a standing mandate. It records preferences, constraints, and, crucially, what happened last time. It is private, invisible to you, unbuyable, and it survives across sessions and across the boundary of any particular query.
Laid side by side, the three barely resemble each other.
| Layer | What writes it | How fast it changes | Who can address it | What it decides |
| Parametric prior (model weights) | Training corpora, months before you read this | Only at retraining; effectively frozen | Nobody, directly | Who gets named with no retrieval |
| Retrieval consensus (per query) | Whatever third parties currently publish about the category | In and out within days; 21.2% source overlap between repeat answers | Marketing, continuously | Whether you make the shortlist at all |
| Delegated memory (per user or per mandate) | Outcomes: what was chosen, what was delivered, what went wrong | Written once, decays slowly, survives sessions | Operations, after the fact | Whether the question is asked again at all |
The Three-Memory Model. Read the last column downwards: shortlisting and re-asking are different events, and no single budget line buys both.
Read the model across, and the awkward pattern appears. The layer you can address fastest is the one that forgets fastest. The layer that remembers longest is the one you cannot address at all. The only lever that connects them is a completed transaction — which is to say, the thing that happens after marketing has finished its job.
Key takeaway
Retrieval consensus decides whether you are considered. Delegated memory decides whether consideration ever happens again. These are separate purchases, and most 2027 plans contain only the first.
3. Retrieval consensus is a rate, not a position
Ranking language has survived into a system it does not describe. Teams still speak of “holding” a slot in an AI answer and “defending” it. The measurement data says something plainer: for almost every brand, presence is a probability, and reporting it as a state is a category error.
The evidence is now large enough to argue from. Beyond the Parse figures already cited, MaxAEO re-ran 1,247 buyer-intent prompts daily for 90 days across eight platforms — 897,840 answers — and found that on any given day roughly 17% of prompts returned a different set of recommended brands than the day before, with the median brand list surviving unchanged for just five days, and three on Perplexity. About a third of that movement is sampling noise; the rest is the platform genuinely changing its mind. GetMentions AI, tracking 530,875 citations across four engines over a week in June 2026, found that 69% of the sources behind a typical answer changed overnight, and that 84% of the sources cited for a given question were used by only one engine.
That last figure deserves a pause, because it quietly kills the idea of a single graph. If four engines answering the same question overwhelmingly cite different sources, there is no shared structure to be inside. There are four separate reassemblies of an argument about your category, running in parallel, agreeing mostly on the very largest names.
The exception that makes the rule useful
Volatility is not uniform, and this is where the honest version of the story gets more interesting than the pitch. Parse found that a typical question had one or two anchor sources that appeared in at least 80% of repeat answers, out of roughly 80 distinct domains cited across the runs. SISTRIX, analysing 82,619 prompts and 1,548,213 snapshots across six countries including the UK over 17 weeks to April 2026, identified a small core of domains whose weekly churn-in is close to zero — cited week after week regardless of what else rotates.
So stability exists. It is just extremely concentrated. One or two sources per question hold something like a position; the remaining seventy-odd hold a sampling probability. The strategic question is not “how do I get in” but “am I trying to become an anchor for a small set of questions, or to raise my appearance rate across a large set?” Those are different content programmes with different economics, and firms routinely fund the second while reporting against the first.
For example: a UK insurance comparison brand appearing in 30% of runs for eleven high-intent prompts is in a materially different position from one appearing in 85% of runs for three. The first is buying lottery tickets across a wide field. The second owns three questions. Only the second survives a bad month.
The measurement consequence is immediate, and it is the least controversial thing in this article. Presence must be sampled, not observed. Parse’s own guidance is that a five-point weekly move in a volatile category such as software is usually noise. Any alerting rule that fires on a single absence is manufacturing false emergencies, and any board slide built on one screenshot is describing a dice roll.
Key takeaway
Stop reporting AI presence as a binary. Report it as an appearance rate across a fixed prompt set, run repeatedly, and decide explicitly whether you are buying anchor status on a few questions or coverage across many.
4. The memory you cannot buy
Now the other layer, which behaves nothing like the first.
Delegated memory is a stored record about a specific buyer, held by the assistant or agent acting for them. It is where “we use these people for washroom consumables” lives, and where “the last order arrived short and they substituted without asking” lives. It is not assembled at query time from the open web. It is written by events.
Two properties make it strategically different from everything marketers are used to.
First, it is largely written without anyone deciding to write it. The 2026 study The Algorithmic Self-Portrait, which analysed 2,050 memory entries drawn from 80 real ChatGPT users, found that 96% of those memories were created unilaterally by the system rather than at the user’s explicit request. The user is not curating a vendor list. The assistant is inferring one, from what it observed.
Second, and this is the part that resists every existing playbook: there is no input. You cannot publish into it, pitch into it, or bid for it. It has no auction, no editorial contact, no schema. The only mechanism that writes your name into a buyer’s delegated memory is having been selected once and then having the resulting experience be worth recording.
In business-to-business categories this layer is further along than the consumer discussion suggests, because it does not depend on a chat assistant remembering anything. A procurement agent operating under a standing mandate holds an approved-supplier list, a set of constraints, and an exception history, and it consults them before it consults anything external. The memory is not a nice-to-have feature of the product; it is the entire point of delegating the task.
Can you get your brand into ChatGPT’s memory?
Not directly, and nobody can sell you access to it. Personal memory is written from a specific user’s interactions and outcomes, not from published content or paid placement. The only route in is to be recommended once through the retrieval layer, be chosen, and then perform well enough that the experience becomes a stored preference rather than a stored complaint.
The commercial upside of that memory is not speculative. Alhena, reporting across 329 brands, found that cross-session memory was associated with roughly four times the conversion rate of memory-less sessions — unsurprising, because a returning buyer whose constraints are already known skips the entire evaluation the retrieval layer exists to run.
Which is exactly the point. When delegated memory holds a workable answer, the open question is never asked. Your competitors’ citation share on that prompt becomes irrelevant to that buyer, because that buyer’s agent stopped consulting the open market. Memory does not help you win the query. It removes the query.
For example: a facilities manager who once asked “who should supply our washroom consumables in the North West” and acted on the answer will not ask it again in month four. Their agent will place a replenishment order. The category’s most visible brand in that month’s AI answers will never be considered, because nothing triggered a consideration.
5. The closing window
If delegated memory removes questions from the open market, then the open market shrinks. That has a straightforward arithmetic consequence which almost nobody prices, and it changes what a citation is worth today.
Work it through with explicit assumptions, all of them adjustable and none of them measured for your category — the point is the shape of the result, not the precision of the inputs.
The Contestable-Share Calculation
Start with the category’s buying occasions per period. Say 10,000 a quarter across UK mid-market buyers.
Estimate the conversion rate to delegation: the share of occasions per period that move from an open question to a memory-held repeat that never re-opens. Say 12% a quarter — conservative in a replenishment category, aggressive in considered one-off purchases.
The contestable pool is what remains open. Quarter by quarter: 10,000, then 8,800, then 7,744, then 6,815, then 5,997.
After four quarters, roughly 40% of category demand has left the open market — while your citation share, measured only against the questions still being asked, can look completely flat.
The inference: the value of winning a citation today is not the sale attached to it. It is an option on a memory write, exercisable only if the fulfilment holds. Price citations accordingly, and expect the option to get more expensive as the pool closes.
The falsifier: if repeat occasions re-open the question at close to the same rate as first purchases — because buyers keep re-checking, or because mandates are re-tendered — the pool does not shrink and this calculation is worthless. Measure your own re-open rate before believing it.
Two things follow that are worth arguing about in a planning meeting. One is that early-category presence is worth more than late-category presence by a factor that has nothing to do with content quality — a structural argument for spending ahead of your comfort, which is unusual, because most such arguments are vendor-manufactured and this one falls out of arithmetic you can check.
The other is more sobering. If citations are options on memory writes, then a citation followed by a poor first order is worse than no citation at all. You have converted a contestable occasion into a locked one, and locked it against yourself.
Key takeaway
Model your category’s re-open rate before you set the 2027 budget. If repeat demand is genuinely leaving the open market, visibility spending is buying options with an expiry, and fulfilment quality is what exercises them.
6. Positive memory is weak. Negative memory is strong.
There is an asymmetry inside delegated memory that decides where the real risk sits, and it follows from how memory systems compress.
A memory store that kept everything would be useless, so these systems summarise: they retain what is distinctive and discard what is routine. A delivery that arrived on the promised day is routine. A delivery that arrived three days late, short two lines, with an unannounced substitution, is distinctive, discrete, dated, and attributable to a named supplier. The first is compressed to nothing. The second survives.
The consequence is that you do not accumulate preference so much as accumulate an absence of reasons to switch. That is a genuinely different asset from the compounding brand equity the graph metaphor implies, and it is managed differently: by eliminating memorable failures rather than by manufacturing memorable successes.
Anyone who has worked through a manual penalty will recognise the shape of this. The recovery process after a manual action is slow, evidential and asymmetric in exactly the same way: the damage is instant and specific, the repair is gradual and has to be demonstrated. What is different here is that there is no reconsideration request. Unlike a bad link profile, where a disavow file gives you a formal mechanism to disown what harmed you, there is no interface through which to contest a record held privately inside a buyer’s agent. You cannot see it, appeal it, or remove it. You can only outlive it.
Memory portability sharpens this further. Assistants have begun importing memory across vendors — Claude can now bring in memories from ChatGPT, Gemini and Grok, a feature Anthropic still labels experimental. The immediate reading is that this lowers switching costs between assistants. The reading that matters for a supplier is that the record travels. A bad entry written in one assistant no longer stays there.
The same logic runs the other way, and it is worth naming as an emerging attack surface rather than a hypothetical one. A private record built from reported outcomes can be poisoned by reported outcomes, and the disciplines developed for defending against negative SEO — monitoring, evidence retention, fast documented correction — transfer more or less directly. The defence is not technical. It is having a clean, timestamped record of what you actually shipped.
How do you measure brand preference in AI assistants?
Not with a single visibility score. Three metrics, owned separately: appearance rate across a repeated prompt set (the contested pool), first-order win rate among buyers who arrived via an assistant recommendation (the audition), and repeat retention among buyers whose agent now orders without re-asking (the contract). A programme can be improving on the first while collapsing on the third, and a blended score hides exactly that.
7. What this looks like with a budget attached
Ashcombe Supply is invented, but the shape is not. Treat every figure as illustrative.
Ashcombe is a UK distributor of facilities and washroom consumables: around £62M turnover, 11,000 SKUs, roughly 4,000 customer sites, heavily weighted to scheduled replenishment rather than one-off purchases. In late 2026 its board approves a £220,000 programme in response to a single observation — an increasing share of enquiries arrive with phrasing lifted verbatim from an assistant, and two large facilities-management groups have told the sales team they now run supplier shortlisting through an internal procurement agent.
Two plans reach the board. Both are competent. Only one of them is complete.
| Line | Plan A: buy the graph | Plan B: buy the audition, then win the contract |
| Earned corroboration | £150k — trade press, comparison sites, category content at volume | £85k — concentrated on eleven questions the firm intends to anchor, not breadth |
| First-order integrity | £0 | £70k — no silent substitutions, honoured delivery windows, machine-resolvable returns |
| Published fulfilment data | £0 | £35k — monthly first-party fill rate, on-time rate and substitution rate, published openly |
| Measurement | £40k — AI visibility monitoring | £30k — appearance rate, first-order win rate, repeat retention, tracked separately |
| Product data hygiene | £30k | £0 — already adequate; the audit said so |
Ashcombe Supply: two £220,000 plans. Illustrative figures.
How the year runs
Month 0. Plan A commissions a content and digital PR programme against 140 category prompts. Plan B runs a smaller diagnostic first: it samples 40 prompts twenty times each, finds an appearance rate of 9%, and separately pulls the last twelve months of first orders from assistant-referred enquiries. That second number is the one nobody had looked at. Of 61 such customers, 21 placed a second order.
Months 1–4. Plan A’s appearance rate climbs from 9% to 31% — a genuine result, and on any AI visibility dashboard a triumph. Plan B’s climbs to 24% across a narrower set, but reaches 71% on four of its eleven target questions. Meanwhile Plan B has stopped the practice of substituting unlike-for-unlike on short lines without notification, and has moved delivery-window confirmation from the despatch note to the order acknowledgement.
Months 5–8. Plan A’s new customers arrive and their first orders behave exactly as Ashcombe’s first orders have always behaved: an 11% substitution rate and a 78% on-time rate. For a human buyer that produces a mild complaint. For a procurement agent it produces a record. By month eight, two of the facilities-management groups have standing instructions that route washroom consumables to an alternative supplier by default. Nobody at Ashcombe is told this, because there was no conversation in which to be told.
Plan B’s substitution rate is 3% and on-time is 94%, and its published monthly fulfilment data has been picked up by two trade publications and a procurement newsletter — which means the operational fix has quietly become earned corroboration, feeding the retrieval layer it was not bought to feed.
Month 12. Plan A reports a 31% appearance rate, 96 assistant-originated first orders, and 29 second orders — a 30% retention rate, roughly where it started. Plan B reports a lower 26% appearance rate, 71 first orders, and 44 second orders: 62% retention. Plan A won more auditions and lost more contracts. It also did something worse than nothing on a slice of its market, converting contestable accounts into accounts locked against it.
The general lesson is not that content spending is wasted. It is that a plan which funds only the retrieval layer is buying options it has no mechanism to exercise, and the failure is invisible on every dashboard it bought.
8. The strongest case against this argument
The serious objection is not that memory is unimportant. It is that memory is far weaker than described here, and that everything is re-derived anyway. Stated properly it goes like this.
The volatility data is on the objector’s side, and it is the best data in this article. If two repeat answers share only a fifth of their sources, and if 17% of prompts change a recommended brand overnight, then whatever is being stored is not producing stable preferences. Agents re-verify hard facts at transaction time — price, availability, lead time — rather than trusting remembered ones, and Gartner has found that 54% of users double-check every AI-supplied fact. Memory is user-editable, now user-inspectable, and increasingly portable, which makes it a weak moat by construction. And the memory literature itself shows entries being revised and expired rather than accumulated. On this reading, “delegated memory” is a thin preference cache and the contestable pool never really closes.
Concede the core of it: on any verifiable fact, nothing is remembered. It is re-checked, every time, by design. A durable preference can therefore never be a durable claim about your price, your stock, or your lead time. That is correct and it disposes of the crude version of the moat argument.
The bound runs four ways.
- The volatility studies sample the contested pool only. Every one of them measures repeated open questions — prompts deliberately re-run. A replenishment that executes without a question being asked generates no answer to diff, so it is definitionally absent from the dataset. High measured churn among contested queries is fully compatible with a shrinking contested pool; the measurement instrument cannot see the mechanism in dispute.
- Stability is concentrated, not absent. The same studies that report heavy churn also report anchors — one or two domains present in 80% of repeat answers, and a core with near-zero weekly churn-in. Averages across the long tail describe the tail, not the head.
- Re-verification asks a narrower question than selection. Checking that the incumbent still qualifies is not the same event as evaluating whether a challenger deserves the slot. A default that survives verification is a default that was never re-tendered, and the challenger never entered the comparison.
- Weak memory still carries the negative record. Even a thin, aggressively expiring store retains the distinctive event, and failures are distinctive by nature. The moat is thinner than vendors claim and the penalty is deeper — which is a reason to take the layer more seriously, not less, and it is the one implication that does not depend on how strong memory turns out to be.
What would actually falsify the argument: engines beginning to sell placement inside personal memory, which would make it a media channel rather than an operations output; or evidence that delegated repeat occasions re-open the supplier question at close to the rate of first purchases, which would keep the contestable pool intact. Both are checkable. The first would be announced. The second you can measure in your own order book this quarter, and it is the cheaper of the two things to find out.
9. What to do on Monday
A literal list, in order, none of it requiring new tooling beyond what a competent programme already runs.
- Pull your re-open rate. Of customers acquired through an assistant recommendation in the last twelve months, what share placed a second order without your team initiating it? That single number tells you whether the argument above applies to your category.
- Split your visibility KPI into three: appearance rate across a fixed prompt set run at least twenty times, first-order win rate, repeat retention. Retire any blended AI visibility score, which averages the audition and the contract into a number that cannot be acted on.
- Decide, explicitly and in writing, whether you are pursuing anchor status on a small set of questions or coverage across many. Fund one. The content programmes are different and the second is usually the default by accident.
- Audit the first-order experience for the failures a memory system would find distinctive: silent substitutions, missed delivery windows, returns requiring a phone call, anything that forces a human escalation. Fix those before increasing citation spend.
- Publish your own fulfilment data monthly — fill rate, on-time rate, substitution rate. It is honest, it is hard for a competitor to copy, and it is corroboration a retrieval layer can actually use.
- Instrument alerting on repeated absence, never on a single miss. Set the threshold above the noise floor for your category, which is higher than instinct suggests.
- Put one operations owner on the retention metric. If nobody outside marketing is accountable for it, the second half of the system has no owner and will not be funded.
None of this replaces the underlying work. Earned corroboration is still what gets you considered, and the mechanics of building it are the mechanics covered across the fundamentals of link building, the tactical range in our guide to link building strategies, and the practical realities of what backlinks actually do. What changes is what that corroboration is for. It is no longer buying a durable position in an index. It is buying an audition in a system that reassembles its opinion of your category several times a week.
The rest of the programme sits where it always has: rigorous competitor backlink analysis to understand who is anchoring the questions you want, guest posting done at a publication standard, selective editorial link insertions where a genuinely relevant page already ranks, and sponsorship placements where a category has real institutions worth being associated with. For firms selling beyond one market, the same three-memory logic applies per country, which is why international link building remains a separate discipline rather than a translation exercise.
Two closing notes on measurement. The benchmark data on link building outcomes remains the sanity check against vendor-supplied lift figures, and the tooling landscape now includes prompt-sampling products that did not exist eighteen months ago — worth evaluating on whether they report rates rather than states. If your team is hiring for this, the role has changed shape enough that the brief for a modern link building specialist should now include measurement design, and the old habit of optimising for featured snippet capture is a reasonable model for what anchoring a question looks like when the answer surface is generative rather than extractive.
The recommendation graph, in the end, is not a place. It is two systems with different clocks, joined by a transaction. Marketing gets you the transaction. Whether it becomes a preference is decided by people whose names are not on the marketing plan.
