Longitudinal Tracking

Longitudinal Tracking: Watching Citations Compound Over Quarters

TL;DR

The comparison the whole market renews on — this quarter’s citation number against last quarter’s — is the one comparison that cannot prove the thing it is used to prove.

A rising citation line can be three different things: compounding (each new placement buys more than the last), accumulation (you simply kept buying) or replacement (the level holds while the composition turns over). They produce an identical chart and imply three different budgets.

In any single series, age, period and cohort effects are exactly collinear. More quarters add observations, not identification. Only a design fixes it.

The fix is a table that general insurance, retail credit and demography already share: the vintage triangle, read three ways.

Cumulative advantage in citation is real but local — it happens inside a claim cluster you already hold, not at brand level.

Only third-party placements carry a fixed date, a fixed text and an external witness, so the earned layer is the only part of an estate that can be tracked longitudinally at all.

1. The comparison the whole market renews on

The generative-engine optimisation market has settled on an accountability standard, and it is a comparison across quarters. The agency-audit checklist published by Acromatico in July 2026 puts it plainly: a retainer that cannot show whether AI answers cite you more than a quarter ago is producing activity, not results. The sales argument sitting on the other side of the same table is also about time. Subscribe PR’s August 2026 cost-benefit piece argues that early movers in a niche hold their citations for months, and calls that compounding effect the core reason to start before a category crowds. Soar Agency publishes a payback curve to the month: early citation wins on long-tail prompts in months three to four, measurable share-of-voice gains on high-value buyer prompts in months five to nine, nine to twelve months in saturated categories.

Every one of those is a claim about how a stock of earned citations behaves as it ages. Every one of them is tested with the same instrument: one number placed next to the number from three months ago. That instrument cannot do the job — and not because it is noisy or because the sample is small. It fails for a structural reason. The quantity it produces is a sum of three effects that no quantity of history will separate.

What is longitudinal citation tracking?

Longitudinal citation tracking follows the same units — the same placements, the same prompt clusters, the same claims — across repeated measurement periods, rather than re-measuring a market snapshot each quarter. The distinction is not pedantic. A snapshot repeated four times is four cross-sections, and four cross-sections answer a different question from one panel. The first tells you what the market looked like on four dates. The second tells you what happened to things you owned. Almost every visibility score sold as a tracking product produces the first and is read as the second.

2. Three things a rising line can be

Accumulation

The stock grows because you kept buying. Each placement contributes roughly what the last one did, and coverage rises approximately in proportion to cumulative spend. This is a perfectly good business. It is also a linear cost with a decay term: stop buying and the line flattens, then falls at whatever rate the existing stock fades. Nothing about it justifies a hockey-stick forecast.

Compounding

The return on the next unit is a function of the stock already held. This is a claim about the marginal placement, not the total: the fifteenth placement in a cluster buys more coverage than the fifth did. It is the only one of the three that justifies concentrating budget, and the only one that makes an early start structurally valuable rather than merely earlier.

Replacement

Gross additions roughly equal gross losses. The level holds while the composition turns over completely. Digital Authority Partners’ 2026 visibility study measured a Citation Retention Rate — CRR, the share of cited URLs holding position between measurement waves — of about 33% over 28 days across five platforms, with 24% of query-and-platform pairs showing zero overlap between the first and third waves. A stable headline sitting on top of that much turnover is not stability. It is a treadmill, and it is funded from a maintenance line, not a growth line.

Key takeaway

Accumulation, compounding and replacement produce the same chart and demand three different budgets: linear spend, concentrated spend, maintenance spend. Deciding which one you are looking at is the entire purpose of tracking longitudinally. If your reporting cannot distinguish them, it is not longitudinal — it is just frequent.

Do AI citations actually compound?

Sometimes, and not at brand level. The evidence for cumulative advantage is strong inside a claim cluster where a brand already holds a position, and weak to absent across a brand’s whole prompt set. That distinction changes where money goes, and it is invisible in any single brand-level visibility score.

3. Why one series cannot tell them apart

Every observation of a placement carries three dates. When it was earned — its cohort. How old it was when you measured it — its age. When the measurement happened — the period. These are not three independent facts. Period minus age equals cohort. Fix any two and the third is determined.

That is the age–period–cohort identification problem, and it has been formally understood for half a century. Norval Glenn’s 1976 paper in the American Sociological Review called the attempt to separate the three statistically a futile quest, and the label has held up: the three effects are exactly collinear, so the design matrix is rank-deficient. This is not a power problem that a fifth or a twentieth quarter of data will solve. Adding quarters adds observations. It does not add identification. Separating the effects requires a restriction you bring from outside the data — an assumption, an experiment or a design.

The diagonal that looks like a decay curve

Here is the practical damage. The standard citation decay chart — current performance plotted against how old each placement is — is a diagonal read of a cohort table. It varies age and cohort simultaneously. If your recent placements are better than your older ones, that diagonal overstates decay, because part of the drop you attribute to ageing is really the difference between a good buyer today and a worse buyer last year. If your recent work is weaker, the same chart hides decay and can even slope upward while every individual cohort fades.

Period effects are not hypothetical here. In the compilation AuthorityTech published in June 2026, the share of AI Overview citations coming from pages ranked in the organic top ten fell from roughly 76% in mid-2025 to 38%, with pages ranked 11 to 100 taking 31.2% and pages beyond rank 100 another 31.0%. A brand whose citations sat mostly on top-ten pages would have logged a steep decline across that window caused entirely by a change in the engine’s sourcing behaviour. Nothing in its own estate aged. Model updates, licensing deals and shifts in how training corpora are assembled all land on the whole market at once, and they arrive in the same column of the same spreadsheet as your own results. Until 2026 the only period effect most search teams had to reason about was an update or a manual action landing on one site — discrete, dated, announced. These are none of those.

What four quarters show that four snapshots cannot

A cross-section can rank you against competitors today. It cannot tell you whether a position was won or inherited, and it cannot price a decision taken last year. A panel yields three quantities that no repeated snapshot contains. The first is the age curve of your own stock, which is what a decay rate is supposed to be. The second is the arrival rate: how long a new placement takes to convert into a citation, which is the only direct read on whether the market is getting easier or harder for you specifically. The third is transitions — which clusters moved from absent to present, which moved back, and in which order — because a net figure of “flat” is produced equally by nothing happening and by two large flows cancelling.

Transitions matter more in this channel than in older ones because the retrieval path itself keeps moving. A placement reached through a conventional crawl in one quarter may be reached by an agent browsing the page on a user’s behalf in the next, and the value of the visit changes with the mechanism. Those are period effects with your name on them, and only a panel lets you see them arrive as a shift affecting every vintage at once rather than as a slow decline you might blame on your own ageing content.

4. The vintage triangle

Three professions independently solved this problem with the same table, and marketers already use it for customers while refusing it for citations.

General insurance calls it a development triangle: rows are accident years, columns are development years, cells are cumulative claims, and loss development factors — LDFs, the ratios between adjacent columns — project the unfinished rows toward their ultimate value. Retail credit calls it vintage analysis, with rows as origination quarters. Demography and epidemiology call it an age–period–cohort model. In 2025, Pittarello, Hiabu and Villegas showed in the North American Actuarial Journal that the chain ladder — the most widely used reserving method in general insurance, accepted by every regulator — is the age-only special case of an APC model borrowed from demography. A profession that has run these triangles for a century turns out to have been assuming all along that there were no period and no cohort effects. That is precisely the assumption buried in the sentence “citations compound”.

What is a citation vintage triangle?

A citation vintage triangle is a table whose rows are the quarter in which a set of placements was earned and whose columns are the age of that set at measurement, with each cell holding the share of that cohort still appearing in at least one cited answer. It converts a single ambiguous line into three separable readings. Below is one programme’s triangle, measured at the end of Q2 2026.

Cohort (quarter earned)At 90 daysAt 180 daysAt 270 daysAt 360 days
Q3 2025  (6 placements)67%50%50%33%
Q4 2025  (9 placements)67%56%44%not yet observed
Q1 2026  (14 placements)71%57%not yet observednot yet observed
Q2 2026  (18 placements)72%not yet observednot yet observednot yet observed

THE VINTAGE TRIANGLE. Shaded cells are the calendar snapshot — every cohort as it stands today. That diagonal is what a quarterly dashboard shows you.

Across a row is the age effect for one cohort, holding the buyer and the vintage constant. This is the only honest decay curve you own: 67, 50, 50, 33.

Down a column compares cohorts at equal age, holding ageing constant. Did the placements bought in Q1 land better at 90 days than those bought in Q3? Here: 67, 67, 71, 72 — a slight, real improvement in buying.

Along the shaded diagonal is the calendar snapshot: 72, 57, 44, 33. It is the steepest line in the table, it is the one on the slide, and it belongs to no cohort. It confounds age with vintage, and in this case it overstates the true fade of the oldest cohort by five points while implying a 180-day retention of 57% that no cohort has actually delivered.

One more borrowing. In reserving the newest row is the least developed: most of its claims are incurred but not reported — IBNR, meaning they exist but have not surfaced. The standard response is not to extrapolate a thin row: the Bornhuetter–Ferguson method anchors an immature year to a prior expectation instead. Your newest cohort has the same property. Placements earned six weeks ago have citations that will exist and have not yet been sampled, and that row is always the one on the current slide. Mark it immature and do not let it set the trend.

Why your own website has no cohorts

A cohort needs two things: a fixed entry date and a fixed unit. Your own pages have neither. You hold the edit rights, and the field’s own maintenance advice guarantees you will use them — the same Digital Authority Partners study recommends that editorial calendars carry a refresh schedule alongside a publish schedule. Every refresh resets the age of the asset. An updated page is a new cohort wearing an old URL, which means the maintenance habit the field most recommends quietly destroys the only variable that could test the field’s own compounding claim.

Third-party placements do not have this problem. A guest article carries a publication date you did not set, text you cannot quietly edit, and a witness who is not you. A newsjacked comment lands on a dated news cycle; a launch-day listing on Product Hunt is pegged to an event. Those are datable assets in the strict sense: the date is a fact about the world rather than a field in your CMS.

Two things to date carefully. An inserted link into an existing article carries two dates — the page’s and the insertion’s — and only the second is its cohort; assigning it the page’s date will make an eight-week-old placement look like a four-year-old survivor. And a durable asset such as a calculator or dataset accrues placements continuously, so the asset is not the cohort. Each placement earned against it is.

The conclusion is structural rather than rhetorical: the earned layer is the only part of an estate that supports longitudinal measurement, and it earns that property from who holds the edit rights rather than from any quality of the content. The same logic explains why dated, signed publication records are worth more to an analyst than to an engine.

5. What the evidence actually says about compounding

Cumulative advantage is real, and it has been identified by experiment rather than inferred from a curve.

Van de Rijt, Kang, Restivo and Patil (PNAS, 2014) intervened in four live systems and conferred small successes at random. On Kickstarter, 70% of projects handed an arbitrary small donation went on to attract further funding, against 39% of untreated projects, and the treated projects drew more than twice as many later donations. Bol, de Vaan and van de Rijt (PNAS, 2018) used a sharper design: applicants who fell just above a research-funding threshold accumulated more than twice the funding of applicants just below it with near-identical review scores, over the following eight years. The design is the whole point. The only way to demonstrate compounding is to hold quality constant and vary early success — which is exactly what a longitudinal series with no design cannot do.

The same literature supplies two qualifications that the retainer pitch tends to skip. First, repeated boosts have diminishing returns: the 2014 team ran a dose test — one donation of 1% of a goal against four donations totalling 4% — and the larger dose did not deliver a proportionally larger effect. The first unit does most of the work. That applies to repeat boosts to a single item; first corroborations across many independent claims are a different purchase with a different curve. Second, where cumulative advantage is strong, results become less predictable from quality rather than more. Salganik, Dodds and Watts (Science, 2006) found that adding social influence to an artificial music market raised both inequality and unpredictability. A genuinely compounding regime is one in which your own effort explains less of your outcome, which is an uncomfortable thing to sell and an important thing to plan around.

What the 2026 market data shows

Semrush’s category-dominance work with Kevin Indig tracked 1,094 US categories monthly inside ChatGPT from January to June 2026, across 50,000-plus brands and 600,000-plus citations. Category owners held first place in 90.4% of month-over-month comparisons — positions, once established, are genuinely sticky. But only 15.2% of categories had a clear owner, 53.7% were unsettled, and in the top half by demand just 11.3% were owned. Where leadership flipped, the median lead had been 1.3 points; where it held, 2.9. So the market contains real persistence and very little of it, concentrated in categories that are already decided.

The shape of that persistence is a step, not a slope. AuthorityTech’s 2026 compilation reports that top-quartile brands by web mentions average 169 AI Overview citations against 14 for the next tier down — a twelvefold gap between adjacent quartiles. MaxAEO’s 90-day longitudinal run (1,247 prompts across eight platforms, 90 snapshots) found roughly 2.3 anchor brands per prompt present in at least 80% of snapshots, with four to nine challengers rotating beneath them. Two populations, two clocks. Compounding, if you have it, is the process of crossing from the second population into the first, and that crossing happens one prompt cluster and one recommendation context at a time — never at brand level, and never evenly.

Order effects: why the sequence of buying changes the result

If returns are increasing inside a cluster and flat across clusters, then the order in which you buy determines where you finish. That is path dependence in the strict sense, and it is testable with arithmetic rather than belief. Take twelve placements and six target clusters, and assume — consistent with the anchor counts in the longitudinal data — that a cluster needs roughly three persistent third-party corroborations before a brand starts being named without prompting. Spread evenly, two per cluster, the programme crosses the threshold nowhere and ends the year with twelve placements and no anchor positions. Concentrated three-deep into four clusters, the same twelve placements produce four anchor positions and two untouched clusters. Identical spend, identical target list, different terminal state — and the difference persists, because anchor positions turn over in only about one month in ten.

The sequencing also compounds sideways. A cluster you genuinely hold tends to be carried into the follow-up turns of a conversation, where a brand mentioned once is available to be reused without a fresh retrieval, so the value of the third placement in a cluster is partly collected in questions you never bought. The honest counter is that concentration is only correct once you know which clusters are worth holding. The first quarter of any programme is exploration and should be spread; the error is not spreading, it is running exploration for eight consecutive quarters because a brand-level average never told anyone when to stop.

Key takeaway

Compounding is local. Depth in clusters where you already appear and breadth to create new clusters are different purchases with different returns, and a brand-level average hides both. Buying one more list placement in a cluster you already hold is not the same trade as buying your first mention in a cluster you do not.

6. Two instruments: the compounding check and the decay estimator

The triangle tells you what happened to each vintage. It does not tell you whether the next placement is worth more than the last. For that you need a measure of the marginal unit, and the cheapest one is time.

THE COMPOUNDING CHECK

Measure the median time from a placement going live to its first appearance in a cited answer. Compute it by cohort, then compute it twice: once across the whole programme, once restricted to the clusters where you already hold a persistent position.

Falling in both — compounding. Early positions are making later ones cheaper. Concentrate budget where the stock already sits.

Falling in aggregate only — composition. You bought easier clusters this quarter, not a stronger position. Do not extrapolate.

Flat in both — accumulation. Fund it as a recurring cost with a decay term and stop forecasting a curve that is not there.

Rising in both — saturation. The clusters you are buying into are already served by incumbents. Rotate targets or change the claim.

How long does an AI citation last?

Persistence medians circulate freely — one 2026 UK tracker running 8,400 prompts across four engines put cross-engine median persistence at 41 days and found 22.7% of citations unchanged past 90 days — but every one of those is a period statistic computed over a mixed population of placements of every vintage and every type. It describes a market window, not your estate. A number that describes your placements requires a cohort you dated when you bought it.

When you do estimate it, you have exactly three options, and they are the three that a national cancer registry has. The cohort approach follows a vintage until it reaches the horizon you care about: accurate, and describing placements bought a year ago. The period approach uses the experience of all cohorts observed during the most recent window: current, and describing a hypothetical placement nobody actually bought. The hybrid approach uses cohort estimates for the early ages where follow-up is complete and period estimates for the later ones. The Office for National Statistics and the National Cancer Registration and Analysis Service publish exactly this construction and document which years used which method — in one childhood survival series, one-year estimates were cohort-based from 2001 to 2016 and hybrid for 2017, while the ten-year estimates were cohort to 2007, period from 2008 to 2016 and hybrid for 2017. Period analysis was introduced for this purpose by Brenner and Gefeller in 1996, precisely because cohort estimates arrive accurate and obsolete.

The transferable rule is not the arithmetic. It is that a statistical authority publishes a single trend line whose estimator changes partway along it and says so in the methodology note. A series whose method changes silently is worthless; a series whose method changes with a footnote is fine. Put the estimator next to the number.

One caution before anyone reports a house decay rate. Decay is not a single parameter. An event-pegged placement fades on the news cycle’s clock, a renewed annual sponsorship listing has a known expiry attached to a contract, and a directory or register entry can sit unchanged for years. A blended decay rate is a weighted average of your own purchasing mix, so it will move whenever the mix moves — which makes it look like a market finding when it is a procurement fact.

7. Worked example: Barmoor Metering

Barmoor Metering is a Newcastle-upon-Tyne company selling smart-metering analytics to UK housing associations, at £6.4M ARR, with a £26,000 quarterly earned-media budget. Its agency reported coverage of the monitored prompt set rising from 9% in Q3 2025 to 14%, then 19%, then 24% by Q2 2026, and presented it as a 167% increase demonstrating compounding.

The placements bought in each of those quarters: 6, then 9, then 14, then 18. Budget had risen alongside. Coverage per placement was flat, which is the arithmetic signature of accumulation — the line rose because the buying rose.

The triangle above is Barmoor’s. Reading across the Q3 2025 row gives the real age curve for a single vintage: 67%, 50%, 50%, 33%. Reading down the 90-day column gives 67, 67, 71, 72 — the buying improved slightly, quarter on quarter. Reading the shaded diagonal, which is what the dashboard had been showing, gives 72, 57, 44, 33: a steeper fall than any cohort actually experienced, and the chart on which the agency’s decay assumption had been built.

The compounding check told the more useful story. Median time-to-first-citation by cohort ran 11 weeks, 10, 9, then 5 — apparently accelerating. Split by cluster, it resolved: across the six prompt clusters where Barmoor had become one of the persistently named brands, the median moved from 9 weeks to 3. Across the other 35, it went from 11 weeks to 10. All of the acceleration was in one sixth of the estate.

In July 2026 the company moved £14,000 of the £26,000 quarterly budget out of breadth and into the six anchor clusters plus two adjacent ones. By the end of Q3 2026, coverage of the monitored set had risen from 24% to 27% on 20% fewer placements, while coverage in the abandoned clusters fell from 12% to 7% — a cost taken deliberately rather than discovered later. The larger change was to the plan: the annual forecast had assumed a compounding curve reaching 40% by mid-2027, and rebuilt on cohort-level numbers it projected 31%. The difference between those two numbers was a hiring decision, and it had been sitting inside a diagonal.

8. Where this argument breaks

The strongest objection is not that the identification problem is unreal, but that most buyers do not need it solved. The argument runs: I do not care why my citations persist. I care whether to keep paying for them. A total-return series answers that question, and the triangle is overhead I will never act on.

That is correct, and worth conceding without hedging. For a single continue-or-stop decision on an entire programme, the aggregate is sufficient and the cohort table earns nothing. Three bounds apply.

  • The moment you reallocate — between clusters, markets, agencies or channels — you are making a comparison the aggregate cannot support, because you are implicitly claiming that one part of the estate is compounding and another is not.
  • The moment you forecast, you are extrapolating one of the three effects, and you have to name which. Extrapolating a period effect as an age effect is the specific error that turns a market-wide improvement in retrieval into a line item called “our compounding asset”.
  • If your programme produces one to three placements a quarter, cohorts that small are noise. Pool two quarters into a single vintage, accept coarser resolution, and do not read a four-cell row as a curve.

A second objection deserves a straight answer: age–period–cohort analysis belongs to demographers with fifty years of data, and you have four quarters. True. With four quarters you cannot estimate the three effects, and you should not try to fit a model. But you can refuse to attribute a movement to one of them, and the refusal costs nothing. The discipline here is negative, not statistical.

The honest negative is bigger than either objection. For most programmes, an honestly built triangle will show accumulation rather than compounding. Accumulation is still worth buying — most marketing is — but it has to be funded as a recurring cost with a decay term rather than sold as a snowball, and some programmes will not survive being described accurately. That is a feature of the instrument, not a fault in it.

9. What to do on Monday

  • Date every placement the day it goes live, in a column you will still have in two years. Nobody can retrospectively date a placement they never dated, which makes this the only item on the list with a deadline.
  • Fix the unit. A placement is the URL plus the text as published. If it is materially rewritten, it leaves its cohort and starts a new one.
  • Build the triangle. Rows are the quarter earned, columns are age in 90-day steps, cells are the share of that cohort appearing in at least one cited answer.
  • Never plot performance against placement age across cohorts. That line is a diagonal, and it is the single most common chart in this category.
  • One variable per claim. To claim decay, compare one cohort across ages. To claim your buying improved, compare cohorts at the same age. Any comparison that moves both is not evidence.
  • Mark the newest row immature and forbid it from setting the trend, exactly as an actuary refuses to extrapolate the freshest accident year.
  • Run the compounding check twice — programme-wide and within held clusters — and make the allocation decision on the second number.
  • Publish the estimator (cohort, period or hybrid) beside every persistence figure you report, internally and to clients.
  • Test the claim with a discontinuity. Take two clusters of comparable difficulty, fund one to depth for two quarters and hold the other flat. That is the cheapest available version of the threshold design, and it is the only evidence of compounding that will survive a finance review.
  • Re-price the engagement to the reading. Depth pricing where the triangle shows compounding, per-arrival pricing where it shows accumulation, and a maintenance line where it shows replacement. A specialist who can defend which of the three is happening is worth considerably more than one who reports the aggregate faster.

None of this makes the measurement easier. It makes it honest, which is a different and more expensive property. A quarterly number can only describe what happened; a design is what lets you say why, and the design has to exist before the first quarter you will want to explain. Everything else in an AI visibility programme can be bought late — the tracking, the tooling, the reporting layer, even the strategy. The dated ledger cannot. It is the one asset in this whole field that has to be started before you know you need it, and the corroboration record built off your own domain is the only part of the estate that will still be datable when someone finally asks the question.

Which is the quiet conclusion of measuring anything in the answer layer. Standards, sampling, variance, denominators, protocols, attribution and valuation all improve the description of a series. None of them identify it. Only a design set down in advance does that, and the field has spent two years buying instruments while leaving the design blank.

Meta title: Longitudinal Citation Tracking: Do AI Citations Compound?
  |  Meta description:
Longitudinal citation tracking, done properly. Why quarter-over-quarter reporting cannot separate compounding from accumulation, and how a vintage triangle fixes it.

Leave a Reply

Your email address will not be published. Required fields are marked *

AI Visibility P&L Previous post The AI Visibility P&L: Putting a Financial Value on Being Cited
AI Advertising Disclosure Rules Next post AI Advertising Disclosure Rules: What the 2027 Regime Means for Sponsored Citations