AI Citation Attribution Pipeline

From Impressions to Influence: Attributing Pipeline to AI Citations

TL;DR

  • Every attribution model in commercial use — last touch, first touch, linear, time decay, Shapley, Markov, even marketing mix modelling — distributes exactly 100% of an outcome. That constraint is an accounting convention, not a property of causation.
  • In a normalised system a zero is not neutral. An unobserved AI exposure does not score nothing and leave the rest alone; its share is transferred to whichever recorded touch sits nearest the sale, usually branded search or direct.
  • So the transfer grows as your citation performance improves — which is why teams report rising direct traffic they cannot explain and then defund the thing that caused it.
  • Epidemiology faced the same problem and abandoned the 100% rule in the 1970s. The attributable fraction gives you a defensible share of outcomes without ever identifying which deal a citation caused.
  • THE CEILING RULE: exposure prevalence among your won deals is a mathematical ceiling on attributable share, never the estimate. Getting from the ceiling to a number requires instrumenting the deals you lost.
  • Earned placements are the only exposure in an AI-influenced buying journey that survives as an artefact you can put in front of a buyer and ask: did you see this?

1. The number on your board slide is an accounting output, not an estimate

Somewhere in your revenue reporting there is a line that says AI-sourced pipeline, and next to it there is a small number. Four per cent. Two per cent. Eleven opportunities. It is the number a CMO gets asked about in every quarterly review, and the number that decides whether the work of getting recommended is funded next year or absorbed into someone else’s remit.

The tooling around that number improved sharply in 2026. Google Analytics 4 added a native AI Assistant channel on 13 May 2026, reaching broad availability on 7 June: visits arriving with a recognised assistant referrer are now tagged with an ai-assistant medium automatically. That is a real improvement on the previous state, in which those sessions scattered across generic Referral rows.

It is also the smaller half of the problem. The native channel does not cover Google AI Overviews or AI Mode, which still report as Organic Search because Google treats them as search clicks, and coverage beyond the recognised assistant list is uneven. The largest hole is structural: an analysis by Attrifast of 41.2 million sessions found that 71% of ChatGPT visits land in GA4 as Direct, carrying no referrer at all, because they originate in native desktop and mobile apps rather than a browser tab. In July 2026 OpenAI shipped a full conversion measurement stack — a pixel, an events API keyed to an Ads Manager account, standard events carrying amount and currency — for advertisers only, with no equivalent on the organic side. Duane Forrester, formerly of Bing Webmaster Tools, called the parallel: this is the 2011 (not provided) settlement being run again, one platform later.

The field has read all of this as a tracking gap, and a tracking gap has an obvious remedy: better instrumentation, custom channel groups, fresher benchmark statistics, self-reported attribution fields. The diagnosis is wrong. Suppose the instrumentation were perfect — every AI exposure logged, every session resolved to a person. The number on your board slide would still be wrong, and wrong in a predictable direction, because of the rule that generates it rather than the data that feeds it.

What is AI citation attribution?

AI citation attribution is the practice of connecting appearances of your brand or pages inside generative answers to downstream revenue outcomes such as opportunities, pipeline and closed-won deals. It fails differently from ordinary channel attribution because the exposure is not an act by the buyer that leaves a record on your property — it is an act by an intermediary, rendered inside a session you cannot observe and often never followed by a click at all.

2. Every attribution model is a conserved-sum rule

Marketing attribution is, as far as I can find, the only measurement system in professional use that requires its answers to sum to one.

This is not a stylistic habit. It is built into the mathematics of every model on the market. Last touch gives 100% to one recorded event. First touch gives 100% to a different recorded event. Linear divides 100% evenly, time decay divides the same 100% with a recency weight, position-based splits it 40/20/40. Algorithmic models are the interesting case, because they are the ones sold as principled. Google’s data-driven attribution and its imitators rest on the Shapley value, defined by Lloyd Shapley in 1953 and characterised by four axioms — and the first of them is literally called efficiency: the payouts to all players must sum exactly to the value of the whole coalition. The third is the null player axiom: a participant who adds nothing to any coalition receives nothing.

Put those two axioms next to a touchpoint your systems never recorded and the result is not a shortfall. It is a transfer. An unobserved participant is formally indistinguishable from a null player, so it receives zero; and because the total is pinned at 100%, the value it would have held is redistributed among the players who were visible. Markov-chain attribution arrives by a different road: the removal effect can only be computed for a state that exists in the transition graph.

Marketing mix modelling, now accessible through Meta’s Robyn and Google’s Meridian, is often proposed as the escape. It fails more gracefully but it fails: MMM decomposes the fitted outcome into channel contributions plus a baseline, and that sums to the fitted total by construction. Anything real but unmodelled lands in the baseline — the part of the business nobody funds.

So the operative fact about a conserved-sum rule is this: a zero is not neutral. It is a donation. The rule does not fail to see your citation. It sees the sale, values the citation at nothing, and hands its share to whoever was standing nearest the money.

THE CREDIT TRANSFER MAP

For each model class: what the rule enforces, where an unrecorded AI exposure’s credit actually goes, and the symptom that shows up in your dashboard when it happens.

ModelWhat the rule enforcesWhere the AI exposure’s credit goesSymptom in the dashboard
Last touchOne recorded event takes the whole outcomeTo branded search or direct — the arrival path the buyer used once already decidedDirect traffic climbing with no campaign to explain it
First touchThe earliest recorded event takes everythingTo the first form fill, which in a 10-month cycle is already latePaid search looks like a demand creator
Linear / time decayA fixed 100% spread across logged touchesDiluted across the recorded set, weighted toward the endEvery channel looks moderately effective; nothing looks decisive
Shapley / data-drivenThe efficiency axiom: payouts sum to the coalition valueReassigned to observed players, because an unlogged touch is a null playerA defensible-looking model that cannot name what changed
Markov removalProbabilities across a closed transition graphNowhere — the state does not exist, so removal is undefinedModel stability that survives real shifts in demand
Marketing mix modellingChannel contributions plus baseline equal the fitted totalInto the baseline: demand the model says would have happened anywayA large, growing, unattributed base nobody has a budget line for
Self-reportedOne free-text or picklist answer per buyerTo whichever touch the buyer can name — a memory, not a causeWord of mouth and Google absorbing everything

Why the transfer grows as your citations improve

The uncomfortable property of a conserved-sum rule is that it does not merely undercount the invisible participant. It undercounts it more as that participant does more work.

Follow the chain. An answer engine names you in a comparison. The buyer does not click; they open a new tab and search your brand. That session lands in branded search or direct. The better your citation coverage gets, the more of these arrivals you generate, and the more the recorded path fills with exactly the two channels a last-touch rule rewards. Your visibility work shows up in your reporting as evidence that branded search is performing. The intuitive management response to a chart like that is to protect the performing channel and trim upstream spend on the content and placements that produced it. That is not a measurement error with a random sign. It is a measurement error with a direction, pointed at your own supply.

KEY TAKEAWAY  Better instrumentation cannot fix a normalisation rule. If the model must allocate exactly 100%, then improving the visibility of one participant necessarily reduces the modelled share of the others — including the participant that caused the arrival in the first place. The question to ask of any attribution output is not how accurate it is, but what it was forced to add up to.

3. The window your model can see opens after the decision

There is a second problem, independent of the first, and together they compound.

6sense’s 2025 Buyer Experience Report, published on 12 November 2025 from a global study of nearly 4,000 B2B buyers across North America, EMEA and APAC, reports that 95% of the time the winning vendor was already on the buyer’s Day One shortlist, and that the pre-contact favourite goes on to win roughly 80% of deals. The same study found buying cycles shortening from 11.3 months to 10.1, and the point of first contact moving from 69% of the journey to 61% — buyers reaching out about six or seven weeks earlier than the year before, while the decision itself stayed where it was.

Read that against what your attribution system observes: touches on instrumented surfaces — a click, a form, a session, an email open — almost all of which happen after first contact. The recorded path is largely the fulfilment of a decision rather than its cause, and a conserved-sum rule applied to that window distributes 100% of the credit inside the residual 5% where shortlists still get reshuffled.

The report contains one more detail worth sitting with. 94% of those buyers used large language models during the process — to summarise reviews, to analyse data, to organise research across long multi-turn sessions — and yet they still averaged 16 interactions with the winning vendor, unchanged from the previous year. The number of touches did not fall. Only their legibility did. This is the same asymmetry that shows up in how agentic browsing devalues the click: the buyer’s work moved to a surface that leaves no trace on yours.

4. The two questions you cannot ask with one instrument

The standard remedy for the dark funnel is self-reported attribution: put a How did you hear about us field on the demo form, cross-reference it against the CRM, and treat the buyer as the sensor of last resort. Practitioners report that this surfaces 30% to 50% of pipeline that digital attribution cannot see, and directionally they are right that it surfaces something. The trouble is what it asks.

How did you hear about us is a causal question: it asks a person to identify, from memory, the cause of their own preference. Richard Nisbett and Timothy Wilson settled the reliability of that class of report in 1977 in Telling More Than We Can Know (Psychological Review 84(3)), across studies in which people confidently named reasons for their choices that could be shown not to be the operative ones. In the best-known, shoppers evaluating four identical pairs of stockings favoured the right-most pair by roughly four to one and denied that position had played any part. They were not lying. Introspective access to the causes of one’s own judgements is poorer than the fluency of the explanation suggests.

A B2B buyer nine months into an evaluation is being asked to do something harder still: to compress a diffuse, multi-person, multi-surface process into one answer for a mandatory field. And notice that the field imposes the same conserved sum on the buyer’s memory that the model imposes on the data. One box, one answer, 100% of the credit. The answer that comes back will be the touch that is easiest to name — a brand, a person, a search engine — which systematically favours anything with a memorable label over anything that arrived as an unattributed sentence inside a generated answer.

There is a different question, and it is the one that can actually be answered: did this happen? Not what made you buy, but whether a given exposure occurred at any point in the evaluation. Exposure is a matter of fact. Credit is a matter of theory. A single field cannot carry both, and every hybrid attribution setup I have seen tries to make it do exactly that.

What should replace “How did you hear about us?”

Replace one causal question with two factual ones, asked in conversation rather than on a form: whether an AI assistant was used at any stage of the evaluation, and which specific published sources the buying group can recall encountering. Both are checkable, both are answerable by more than one person in the group, and neither asks the buyer to perform an attribution they are not equipped to perform.

5. What epidemiology did when it faced this exact problem

Marketing is not the first field to need a defensible share of outcomes for a cause it cannot trace at the individual level. Epidemiology got there first, and its answer is the most useful piece of borrowed machinery available to anyone whose job title now includes this work, trying to value citations in generative answers.

The structure is identical. An outcome is observed; exposure is diffuse, historical and partly unrecorded. Nobody can say which cigarette caused which tumour, and no instrumentation will ever say it. What the field built instead was the attributable fraction — Morton Levin’s 1953 population formula, refined by Olli Miettinen in 1974 into the version that works from exposure prevalence among cases rather than in the population at large.

AF among cases  =  Pc  ×  (RR − 1) / RR

where Pc is the proportion of cases (won deals) in which the exposure was present, and RR is the risk ratio between exposed and unexposed groups. Miettinen, 1974.

The move that makes this work is the one marketing never made. Two years after Miettinen, Kenneth Rothman published Causes (American Journal of Epidemiology, 1976), setting out the sufficient-component-cause model: an outcome arrives through a mechanism assembled from several component causes, any one of which, removed, would have prevented it by that route. The direct consequence is that attributable fractions for different causes can and routinely do sum to more than 100%, with no upper limit in principle.

This is not a fringe position. Rothman criticised Doll and Peto in 1986 precisely for constraining their attributable fractions for cancer causes to stay under 100%; Peto’s later revision raised the ceiling to 200%. Rowe and colleagues devoted a 2004 paper in the American Journal of Preventive Medicine to explaining why the sum exceeds one — the same case can be prevented in more than one way, so it is legitimately counted more than once. Nancy Krieger put the general case in the title of a 2017 American Journal of Public Health paper: the fallacy of treating causes as if they sum to 100%.

Marketing attribution is that fallacy, formalised, automated and sold by subscription. Epidemiology gave up the hundred per cent in the 1970s and received something considerably more useful in exchange: a number you can defend in front of someone whose job is to disbelieve it.

KEY TAKEAWAY  Attribution asks: of this outcome, what share belongs to each channel? The attributable fraction asks a different question: what share of these outcomes would not have occurred by this pathway had the exposure been absent? The second question can be answered without identifying a single individual deal, and the answers to it are not required to add up.

6. The ceiling rule

Look again at the formula. Because RR is at least 1 whenever exposure is associated with the outcome, the factor (RR − 1) / RR always falls between 0 and 1, approaching 1 only as RR approaches infinity. Multiply anything by a number less than one and you get something smaller. So:

THE CEILING RULE

The share of your won deals in which an AI answer appeared is a mathematical ceiling on the share of those deals attributable to it. It is never the estimate.

AF < Pc for every finite RR. Equality would require that no unexposed buyer ever bought from you — that AI exposure was strictly necessary for the sale. Nobody believes that, and no vendor deck states it, yet the prevalence is routinely reported as though it were the attributable share.

Three numbers get you from ceiling to estimate: exposure prevalence among wins, exposure prevalence among losses, and the ratio between them. Two of the three are not in your CRM today.

This single inequality sorts the market’s claims into two piles, much as the older argument over citations versus backlinks turned on what was being counted. When a report says AI influenced 43% of pipeline, it has almost certainly measured a prevalence and printed it as a share — the ceiling, reported as the estimate. When a report says AI sourced 4% of pipeline, it has run a conserved-sum allocation over the visible arrival path — a transfer, reported as the estimate. Both are wrong in opposite directions, and the honest answer sits between them.

It also explains why published conversion multiples for AI traffic vary so wildly that they cannot all be describing the same phenomenon. Ahrefs found AI assistants driving about 0.5% of its visitors but 12.1% of its signups, a factor of roughly 23; a Microsoft Clarity study of 1,200 sites put the gap at eleven times. Meanwhile the largest peer-reviewed study of the question — Kaiser and Schulze in Marketing Science — 973 websites, $20 billion in combined revenue, 50,000 ChatGPT-referred transactions against 164 million from other channels — found organic LLM traffic converting below every traditional channel except paid social. The authors name the reason in the paper: the data is last-click. Every one of these figures is conditioned on arrival, and arrival is precisely the event an answer engine is least likely to produce.

7. The exposure ledger: collecting the two numbers you are missing

An attributable fraction needs a case group and a comparison group. Your case group already exists — the deals you won. The comparison group is the one nobody instruments, and building it is the whole job.

THE EXPOSURE LEDGER

Five design rules for collecting exposure rather than credit.

1. Ask factually, not causally. Was an AI assistant used at any point in evaluating this category? — not what made you choose us. Facts survive recall; causes do not.

2. Ask by buying group, not by lead. One person’s answer is a sample of one from a committee. Two or three answers per account change the estimate materially.

3. Ask about the research period, not the arrival. The exposure you care about happened before the form fill, and often before the shortlist.

4. Use aided recall for documents, unaided for tools. Nobody remembers a source name unprompted. Show a short list of named, dated third-party pieces and ask which ones they encountered.

5. Ask the same questions of every loss and no-decision. This is the rule that produces the denominator, and it is the one that will not survive contact with a sales team unless someone senior insists.

Why your lost deals are the most valuable rows in the CRM

Attribution has never instrumented a loss, for a reason that has nothing to do with technology: credit for a loss has no claimant. No channel owner volunteers to report on the deals their channel touched and failed to win. The discipline evolved without a control group, and every number it produces is computed on the survivors.

That omission is exactly what makes the ceiling unbreachable. Without exposure prevalence among the deals you did not win, there is no RR, and without an RR the only number available is the prevalence itself. For example: a firm finds AI exposure in 55% of its wins and announces that AI influences over half its pipeline. Ask the same question of its losses, find 48%, and the picture changes completely — the exposure is nearly as common among buyers who chose a competitor, and the attributable share collapses to something in the low teens. Same data, one extra question, an answer that is roughly a quarter of the headline.

8. Worked example: Alderbeck Treasury Systems

Alderbeck Treasury Systems sells cash-management software to mid-market finance teams from an office in Manchester: £6.2M ARR, a median contract value of £155,000, a nine-to-twelve-month sales cycle, and a marketing team of four reporting to a CFO who had begun asking pointed questions about the content budget.

The January board pack carried a line reading AI-sourced pipeline: 4% — three closed-won deals in which someone had typed ChatGPT into the How did you hear about us field, £236,000 against £5.89M of new ARR from 38 wins. The head of marketing did not believe it, could not say why, and had lost the argument twice.

The rebuild ran over two quarters.

  1. February. The form field was retired. Two picklist questions went into the deal-review template, asked of every opportunity at close regardless of outcome: whether an AI assistant had been used in evaluating the category, and which of eleven named published pieces — six third-party, five owned — the buying group recalled seeing.
  2. March to June. Reps recorded answers for 38 wins and 71 losses and no-decisions. Multi-contact accounts were coded as exposed if any member of the buying group reported exposure.
  3. July. The analysis, which took an afternoon in a spreadsheet.

Exposure appeared in 21 of the 38 wins — a prevalence of 55.3%, and the ceiling on anything Alderbeck could claim. The arithmetic is simple enough to put behind a calculator. It also appeared in 34 of the 71 losses, 47.9%. The odds ratio between the two is 1.34, giving an attributable fraction among cases of 0.553 × (0.34 / 1.34), or 14.2%.

THE THREE NUMBERS, SIDE BY SIDE

Reported by the CRM (conserved-sum allocation): 4% — £236,000

Defensible attributable fraction: 14.2% — £836,000

Ceiling (exposure prevalence among wins): 55.3% — £3.26M

The reported figure understated the defensible one by a factor of 3.5. The prevalence overstated it by a factor of 3.9. Neither error was a tracking failure.

The finding that changed the budget was in the second question rather than the first. Of the 21 exposed wins, 14 buying groups could name at least one third-party piece they had encountered — a trade-title feature won off a timely news hook, a supplier-comparison page on an industry body’s site, two practitioner posts. Four named an Alderbeck-owned page; nothing else cleared five mentions. The owned estate was doing its job at the validation stage and almost none of the work at the selection stage, which is where the 6sense shortlist evidence says the deal is decided.

Alderbeck moved roughly a third of its content-production budget into earning placements on the four outlets that had shown up unprompted in buyer answers — the only exposures in the evaluation they could name, date and buy more of. The January reporting could never have produced that reallocation, because it had already given the credit to branded search.

9. Where this leaves earned links

Third-party corroboration — the oldest asset in the discipline — is the single most systematically undercredited input in any conserved-sum model, and it fails three separate tests at once. It is off-domain, so you cannot tag it. It is upstream, so it sits outside the observable window that opens at first contact. And it is mediated — its effect reaches you through an engine and then through a branded search, so even when it works perfectly the arrival wears someone else’s label. Three strikes in a system that pays out for proximity to the sale.

And then the inversion. The exposure you cannot measure is the answer itself: ephemeral, unrepeatable, different for two people asking the same question in the same hour, gone the moment the session closes. You cannot show a buyer the answer they saw. You can show them the article. An earned placement persists as an artefact with a URL, a date, a publisher and a headline — which makes it the only exposure in an AI-influenced buying journey that supports aided recall, and aided recall is the difference between a survey question people can answer and one they cannot.

That has a direct procurement consequence, and it is the practical payload of this whole argument. If your exposure instrument is built from named documents, then the specification for a placement stops being a domain-authority threshold and becomes a question about recognisability: could you put this in front of a finance director nine months later and get a yes? A trade title your buyers read clears that bar; a guest post on a general marketing blog does not, whatever its metrics say. The test quietly demotes bought edits inside old articles — nobody recalls seeing a link that was inserted into a page they never read — and promotes the formats that arrive with a masthead attached, including sponsorship placements, named research collaborations and the rest of a deliberately chosen strategy set.

10. The strongest objection

Here is the best version of the counter-argument. Nobody who builds attribution models believes they are causal. They are budget-splitting devices. Budgets are conserved — you have one pound to allocate and it goes to exactly one place — so a rule that distributes exactly 100% is not a philosophical error, it is a match between the tool and the decision. The efficiency axiom is in the Shapley value because Shapley was solving an allocation problem, not a causal one. And the complaint that AI citations get too little credit runs into an awkward fact: you cannot buy a citation the way you buy a click. Adding an unpurchasable participant to a spending split is arguably my category error, not theirs.

Most of that survives. A conserved-sum rule is the right tool when every participant is a channel you can buy more of, and I would not replace last-touch reporting for paid search with an epidemiological framework. The attributable fraction has real costs of its own: it needs a control group, a stable exposure definition held constant all year, and enough deals to matter. Below about thirty wins a year the confidence interval swallows the estimate — and splitting the sample across several markets makes that worse, not better.

The boundary is precise, though. A normalised split is harmless while the mix is closed. The moment a participant carrying material weight is neither purchasable nor observable, the split stops being a tie-breaker and becomes a transfer — and the transfer runs away from the input that produced the credit and toward the channel that merely recorded the arrival. A rule that is merely imprecise costs you accuracy. A rule that is biased in a known direction costs you the budget line that was working.

One honest concession to finish. Run this properly and your attributable fraction will sometimes come out lower than the number your CRM was already reporting, particularly in categories where nearly everyone is exposed and exposure therefore discriminates between nothing. Publish that result too. An instrument that can only revise your budget upward is a sales tool, not a measurement.

11. What to do on Monday

  • Find out what your attribution model is forced to add up to. If the answer is 100%, you now know where an unrecorded exposure’s share went — write down which channel received it.
  • Pull your AI Assistant channel in GA4, then treat it as a floor rather than a measurement. With 71% of ChatGPT sessions landing in Direct, the visible slice is a sample of unknown selectivity.
  • Retire the single How did you hear about us field, or demote it to a curiosity. It is asking a causal question of a witness who does not have the answer.
  • Add two picklist questions to the deal-review template: was an AI assistant used, and which named sources does the buying group recall. Fix the wording now, because you cannot compare quarters across a changed question.
  • Ask both questions of every loss and no-decision. This is the only step that turns a ceiling into an estimate, and the only one that will meet resistance.
  • Compute the attributable fraction after two full quarters of collection, not before. Report it with the ceiling and the CRM number next to it, all three labelled.
  • Rank your earned placements — including launch-surface listings — by whether a buyer could plausibly recall them nine months later, and let that ranking, rather than a third-party authority metric, drive next quarter’s prospecting.
  • Keep a dated register of every placement earned, and note which of them are reachable by the engines at all. It is the only census-grade record in the whole programme and it takes a spreadsheet.

The citation layer has not made attribution harder. It has revealed how much of attribution was always a convention — an agreement to pretend credit is a conserved quantity because the arithmetic is tidier that way. That held while every participant was a channel with a tag on it. It does not hold now, and the tell is that the more your answer-layer visibility improves, the smaller your reporting says it is.

Leave a Reply

Your email address will not be published. Required fields are marked *

Benchmarking Citation Share Previous post Benchmarking Citation Share vs Competitors: A Reproducible Method
AI Visibility P&L Next post The AI Visibility P&L: Putting a Financial Value on Being Cited