TL;DR
The deadline that justified the moat was cancelled. Google kept third-party cookies in April 2025 and retired most of the Privacy Sandbox in October 2025. The customer-data land grab was sold against a deprecation that never happened.
Agent commerce did not fail on customer knowledge. ChatGPT’s Instant Checkout was scaled back in March 2026 with roughly 30 Shopify merchants live, because inventory, shipping and delivery facts were wrong — not because the assistant knew too little about the buyer.
Recognition is the thing agents break. Stripped referrers, no JavaScript execution, no cookie, no banner interaction: the visitor you most want to identify is the one your stack cannot see.
The asset that survives is permission, not the pile. Value has moved from data you use on the buyer to data you can release to whoever is making the decision — and almost no consent architecture in Britain is built to permit anything.
The blocker is internal. UK law got more permissive in 2026, not less. What stops most companies publishing their own numbers is a retention schedule, a privacy notice and one clause in a customer contract.
The link-building payoff is a statistic nobody else can compute. Operational data you are permitted to release becomes a recurring, citable figure — an asset most businesses already own and have never been allowed to say out loud.
1. The moat was built for a deadline that never arrived
For six years, every first-party data programme in Britain was sold with the same slide: third-party cookies are going away, so build a direct relationship with your customer before the lights go out. Then the lights stayed on. In April 2025 Google confirmed it would not roll out a standalone prompt for third-party cookies in Chrome and would keep its existing controls. In October 2025 it retired most of the remaining Privacy Sandbox technologies — Topics, Protected Audience, Attribution Reporting and around seven others — citing low adoption, and the Competition and Markets Authority released Google from its Privacy Sandbox commitments and closed a four-year investigation.
That does not mean the work was wasted: Safari, Firefox and Brave still block third-party cookies by default, leaving roughly a fifth of global traffic cookieless whatever Chrome does. But the reason most teams were given was wrong, and the strategy underneath it deserves re-examining — accumulate customer data, because knowing more about the buyer than your competitors do is a durable edge. The channel that premise was supposed to protect you in has now arrived, and it behaves nothing like the one it was designed for.
What the biggest agent-commerce experiment actually broke on
OpenAI and Stripe launched the Agentic Commerce Protocol — an open standard for agent-run checkout — on 29 September 2025, alongside Instant Checkout inside ChatGPT. Six months later it was gone. In March 2026 OpenAI scaled the feature back to product discovery with a redirect to the merchant’s own site. Forrester principal analyst Emily Pfeiffer counted roughly 30 Shopify merchants live on it by February 2026. The reporting around the retreat was consistent about why: onboarding was heavy, and inventory, shipping cost and delivery information were frequently wrong because OpenAI was scraping retail sites rather than being handed accurate feeds.
The most heavily resourced agent-commerce product of the cycle did not fail because the assistant knew too little about the shopper. It failed because merchants could not hand over accurate facts about themselves fast enough for a machine to act on them. The scarce input was never the customer profile. It was the merchant’s own operational truth, in a form somebody else could read.
What is first-party data? It is information you collect directly from your own customers and systems — orders, sessions, support tickets, returns — rather than data bought from a broker or inferred by a third-party tracker. The category is defined by who collected it, which is why it gets confused with the question that matters far more in 2027: what are you allowed to do with it?
The traffic, meanwhile, is real. Adobe Analytics, across more than a trillion visits to US retail sites, put AI-referred retail traffic up 138% year on year in May 2026. In March 2026 those visitors converted 42% better than non-AI traffic — a complete reversal from twelve months earlier, when they converted at roughly half the rate. Keep the denominator honest, though: Contentsquare, benchmarking 99 billion sessions, put AI referrals at around 0.2% of all visits. The channel is small, compounding fast, and the visitors it sends are pre-qualified by somebody else’s ranking before they reach you. That is the structural shift behind the zero-click traffic model — the evaluation happens somewhere you do not own, and you inherit whatever conclusion it reached.
2. The recognition collapse: what an agent leaves behind
Every personalisation strategy rests on recognition: you know who this is, so you show them something different. Agent-mediated traffic dismantles that chain link by link, and it does so quietly, because what breaks does not throw an error.
Google’s AI Mode strips referrers by design, so even human clicks arrive with no origin attached and land in direct traffic. When an agent fetches rather than browses, no JavaScript runs, so no tag fires: no pageview, no session, no conversion event, and no first-party cookie, which means the client ID that ties a visitor to a history is never created. And because a fetch never touches a consent banner, a consent-gated stack would suppress the tags even if the script had run. Anyone who has audited JavaScript crawl behaviour for backlinks already knows the shape of this problem: the machine that matters does not execute your front end.
| Signal you rely on | Direct visit | AI-referred click | Agent fetch or API call |
| Traffic source | Known | Often stripped — lands as direct | Only in server logs, by user-agent |
| JavaScript execution | Yes | Yes, if a human browser follows | No — tags never fire |
| First-party cookie or session ID | Set | Set on arrival, no prior history | Never set |
| Consent decision | Recorded | Recorded for that device only | Nobody clicks the banner |
| Identity at checkout | Account or email | Account, if they log in | Delegated token, scoped to the order |
| Behavioural history you can act on | Full | None at the moment of decision | None |
Two things follow. The first is measurement: a growing share of your highest-intent activity is filed in the junk drawer of attribution, indistinguishable from bookmark visits. That is solvable — server-side logging and user-agent classification recover most of it, and a SERP-less visibility audit is the right frame for reading the logs. The second is not solvable, because it is not a bug. The buyer arrives already qualified: the comparison, the shortlisting and most of the objection-handling happened inside a conversation you never saw, across a chain of follow-up questions that behaves nothing like a keyword — the pattern documented in work on multi-turn query chains. Your personalisation engine gets its turn after the decision has effectively been made.
The pile is also getting less representative
There is a slower problem underneath. Your first-party dataset is not a census of your market; it is a census of the people who both reached you directly and agreed to be measured. A CHI 2025 study of 254,148 sites found that only 15% of the top 10,000 EU websites ran a minimally compliant cookie banner at all, and Advance Metrics, across more than 1.2 million visitors, found 25.4% accepting everything on first click while 33.6% ignored the banner entirely. Now add agent traffic, which never consents to anything: Gartner expects 40% of enterprise applications to embed a task-specific AI agent by the end of 2026, up from under 5% in 2025. The rise of AI browsers makes that concrete — the same human, running the same query, becomes invisible to you depending on which client they opened.
Key takeaway
Personalisation depends on recognition, and recognition is exactly what agent-mediated demand removes. The moat is not being attacked. It is being routed around.
3. Two kinds of first-party data — and only one of them still compounds
Split your customer data by direction of travel rather than by source, and the picture clarifies immediately.
Inward data is data you use on the buyer: segments, propensity scores, retargeting audiences, lifecycle emails, dynamic pricing, on-site personalisation. Its value depends on a rendering surface you control and an identity you can resolve, and agent commerce degrades both.
Outward data is data you hand to whoever is making the decision: stock position, real delivery windows, size and fit behaviour, compatibility, eligibility rules, returns experience, failure rates. Its value depends on being accurate, current and legible to a third party — the exact category the Instant Checkout retreat proved was missing. Some of it is product data; the interesting part is derived from customers, which is why it collides with consent.
Take any dataset you are proud of and ask two questions: could a competitor obtain the same thing, and are you permitted to release it?
| You are permitted to release it | You are not permitted to release it | |
| No one else holds it | Releasable advantage. The only cell that compounds. Your fit data, failure rates, real lead times, resolution patterns — publishable, checkable, and impossible to second-source. | Locked vault. The pile most teams call a moat. Rich, genuinely exclusive, and legally inert. It can inform your own decisions and nothing else. |
| Others hold it too | Commodity proof. Worth publishing for completeness — specs, prices, availability — but it wins nothing, because the engine can get it from four other places. | Dead weight. Storage cost and breach surface. Audit it, minimise it, and stop budgeting against it as if it were an asset. |
The rule that falls out is blunt enough to put on a wall: data you cannot show is data you do not have. Not in the sense that it is worthless, but in the sense that it cannot influence a decision taken by someone else, and most decisions now are.
What counts as outward data? Anything about your own operations that a buyer or an assistant would otherwise guess at, verify elsewhere, or take on trust. If a competitor could only match it by running your business for a year, it is worth releasing; if they could match it by reading your specification sheet, it is commodity proof.
This is where the machine-readable side of the work pays off rather than the marketing side: released data has to be fetchable and structured — the discipline behind AI-readable API feeds — because an assistant running an extended research pass will pull from whatever it can parse in the time it has. Anyone who has watched a deep research mode work through a shortlist has seen the pattern: the source that answers the checkable question gets used, and the one behind a form fill gets skipped.
4. Permission is the constraint — and the regulator is not the one blocking you
Here is the part most marketing teams have backwards. British law became more permissive about statistical reuse in 2026, not less. The Data (Use and Access) Act 2025 — the UK’s post-Brexit rewrite of its data rules — received Royal Assent on 19 June 2025. Section 71 rewrote the purpose-limitation rule in Article 5(1)(b) of the UK GDPR and inserted a new Article 8A on further processing, with new Articles 84B to 84D creating a statutory presumption that further processing for statistical purposes is compatible with the purpose you collected the data for. Section 67 widened “scientific research” to cover privately funded commercial work. A key condition in Article 84B is that the data is converted into non-identifiable information.
In other words, computing a statistic from your own customer records and publishing the aggregate now has a statutory route where it once required a careful balancing assessment. The ICO has been consistent that the Article 89(1) safeguards still apply and that a presumption of compatibility excuses neither transparency nor a data protection impact assessment.
So the door opened. Almost nobody walked through it, because the things that actually stop you are internal. There are four, and only the first is a legal question.
| Precondition | What it actually requires | Where it usually breaks | Who owns the fix |
| Non-identifying output | Aggregates where nobody can be singled out or linked across records, tested at the granularity you intend to publish | Small cells: one figure covering four customers in one postcode is not anonymous | Data protection lead |
| Transparency at collection | A privacy notice that names statistical use and external publication as purposes, before the data arrives | Notices list fulfilment, service improvement and marketing, and stop there | Legal, with marketing input |
| Retention long enough to matter | A de-identified extract kept for the length of the series you want to publish | A 12- or 13-month purge that quietly makes any trend impossible | Engineering and data |
| No contract forbidding it | Customer, marketplace and platform terms that permit publishing aggregate insight | Processor-role data, confidentiality clauses, marketplace terms | Commercial and sales |
The cookie exception that forbids the thing you want to do
There is a trap inside the reform most UK sites have already adopted. Section 112 inserted a new Schedule A1 into the Privacy and Electronic Communications Regulations — the UK’s cookie rules — and paragraph 5 removes the consent requirement for storage or access used solely to collect statistical information about how a service or website is used with a view to improving it. Those provisions were commenced on 5 February 2026, the ICO finalised its supporting guidance on 29 April 2026, and the maximum penalty under the cookie rule rose from £500,000 to £17.5 million or 4% of worldwide turnover.
The exception is narrow: the data must be used solely for that improvement purpose and must not be shared with anyone else except to assist with those improvements. Analytics collected under the free lane is, by construction, unpublishable. If you plan to release a figure derived from site behaviour, you cannot ride the exception to collect it — you need consent, or a different lawful route, and you need the notice to say so.
The European position rhymes. The Digital Omnibus proposal published in November 2025 moves cookie rules into the GDPR through new Articles 88a and 88b. Article 88a(3) sets out a closed list of low-risk purposes that need no consent, including aggregated audience measurement carried out by the provider solely for its own online service and its own internal use. The Council dropped Article 88b in June 2026 and nothing is likely to be in force before late 2027 — a timeline anyone running campaigns across European markets should plan around rather than wait for.
The asymmetry worth internalising
Both regimes have now built a fast lane for measuring yourself and no lane at all for telling anyone what you found. The permission to release has to be created deliberately and up front, and it is the one thing no compliance vendor sells you, because every product in the category is built to restrict.
The processor trap, and a contractual window that is closing
For B2B software companies the blocker is usually not the privacy notice at all. If you hold data on behalf of your clients, you are a processor, and the data is not yours to publish. Whether you can turn it into a benchmark depends on one clause in your data processing agreement — the sentence that grants you rights over de-identified and aggregated data. The IAPP’s 2025 vendor management survey put standalone DPAs in more than 92% of enterprise SaaS contracts, up from around 60% in 2020.
What the clause says varies enormously. Vendor-favourable drafting permits processing de-identified aggregate data for “service optimisation, benchmarking and product development”; the more generous versions add “and external publications”. Those three words are the difference between an annual industry report and a compliance incident. And buyer-side negotiation playbooks published through 2026 now explicitly target the clause, advising customers to refuse any aggregated-data escape hatch they cannot see. If you intend to publish benchmarks from client data, the window in which that clause is easy to sign is closing, and it will close first with your largest customers.
5. A worked example: twelve weeks at a Cumbrian outdoor retailer
Thirlmere Outdoor is a composite rather than one company: a Kendal-based retailer of walking boots and waterproofs, £14m turnover, 2,400 SKUs, two shops, a website and a listing on a large outdoor marketplace. Footwear returns run at 31%, and 62% of those are size exchanges rather than faults. Six years of order lines record which boot, in which size, came back and what it was swapped for. Nobody else in Britain holds that dataset for those styles.
Week 0, the audit. Four blockers, none of them a regulator. The retention schedule purged order-line detail at 13 months. The privacy notice named order fulfilment, service improvement and marketing with consent, and said nothing about statistics or publication. Site analytics had moved onto the new PECR statistical exception in February 2026, making the behavioural data unpublishable by construction. And the marketplace agreement treated aggregate performance data as confidential, covering 22% of the order history.
Weeks 1 to 4, the permission work. The privacy notice was rewritten to name statistical analysis and the publication of non-identifying aggregates as purposes. A de-identified extract of order lines — style, size purchased, size exchanged, month, no customer key — was carved out of the retention schedule and given a five-year life, while raw order data kept its 13-month clock. An identifiability assessment set a minimum cell size of 50 orders before any figure could be published, following the ICO’s March 2025 anonymisation guidance. The marketplace-sourced portion was excluded rather than argued about.
Weeks 5 to 8, publication. Forty fit pages went up, one per boot style, each carrying a single dated, checkable sentence: this style runs half a size small, 58% of exchanges on it were an increase of one half size, n = 412, twenty-four months to June 2026. The pages were plain, server-rendered and linked from the product pages — an unglamorous piece of technical SEO work in service of links, not a content campaign.
Weeks 9 to 12, distribution. The same figures went into the product feed and the structured data, and twelve emails went out — to walking clubs, two national outdoor titles, a mountaineering council, and four forum moderators who had spent years answering “does this boot run small?” by hand.
What happened. By week 12 the forty pages had earned nine referring domains: two outdoor magazines, a walking club federation, five forums and blogs, and a university outdoor society reading list. Over the following quarter, returns on those forty styles fell from 31% to 26.4%, worth roughly £71,000 in return-handling and re-stocking — a bigger number than the links produced. AI-referred sessions to those pages grew 3.1x against 1.6x for the rest of the catalogue, though a simultaneous feed fix confounds that comparison and it was reported internally as such.
What did not happen. Nothing moved on “walking boots”, the head term, and no link arrived from a high-authority domain. And the flagship idea — an annual British boot-fit report — could only cover twenty-four months, because everything before July 2024 had already been deleted under the old schedule. The first three-year series lands in 2028. That is the cost of the retention clause stated in the only currency that matters: not a fine, but four years of waiting for a dataset they thought they already owned.
The same shape in other sectors
The lesson generalises past retail. A recruitment platform knows real time-to-hire by role; a payments business knows genuine chargeback rates by sector; an installer knows how long the part actually takes. In each case the fact is a by-product of operations, nobody else can compute it, and whether it can ever be said out loud was settled by paperwork written before anyone thought to ask.
6. The strongest objection: you are telling me to give away my only edge
The serious counter is not about privacy. It goes like this. Operational data is a differentiator precisely because rivals do not have it. Publish your fit table and every competitor reads it by Friday, adjusts their sizing guidance, and you have paid for their product research. Worse, you have handed the assistant the information it needs to confidently recommend someone else’s boot in the size you established. Meanwhile short factual strings are exactly what models state without attribution, so you are giving away a real asset for a citation that may never carry your name.
That objection is correct on its facts. Disclosure is irreversible, competitors do read, and a model will often state a number without saying where it came from. Three things still favour releasing, and one carve-out favours not.
- The information gets published either way, worse. Fit, reliability and delivery facts already circulate — in reviews, forums, returns threads and competitor comparison pages. If you do not publish the accurate version, the assistant assembles one from third-party guesses, and you inherit that version anyway with none of the control. This is the same dynamic as competitor misinformation in AI answers: the absence of an authoritative statement is not neutrality, it is a vacuum somebody else fills.
- The asymmetry is renewal, not the number. A competitor can copy this quarter’s figure. They cannot copy next quarter’s, because they do not have the order lines. A one-off statistic is a gift; a maintained series is a subscription only you can renew.
- The comparison is not fact-page versus long guide. It is fact-page versus no page. The audience for a specific operational number — a fit ratio, a lead time, a failure rate — is someone who intends to verify it, and verifiers follow links even when the model that surfaced the claim did not attribute it.
The carve-out is real. If your advantage genuinely is the secrecy of an operating parameter — a pricing model, supplier terms, a yield curve, an underwriting rule — do not publish it. The test is whether the fact helps a buyer choose correctly or helps a rival price against you. Fit data helps the buyer. Margin data helps the rival.
A shorter objection: “we will just update the privacy notice”
You can, and you should, and it will not fix the past. A notice updated today governs data collected from today: nothing in the DUAA reforms makes a purpose retroactive, and no notice restores records your retention job has already deleted. Every month you delay does not add a month to the project, it removes a month from the eventual dataset. The penalty for the wrong wording is not a fine. It is that your first original statistic is a year further away than it needed to be.
7. What this actually buys in links: the permission-gated statistic
Original research earns links. That is not new, and it is the most over-recommended tactic in the discipline — usually as a survey, which is expensive, easy to fake and impossible to verify. The version that follows from this article differs in three ways.
It is a census, not a sample: you are not asking 500 people what they think, you are reporting what every order in your system did. It is checkable, because it names an n, a period and a definition. And it is gated by permission rather than budget, which is why so few competitors publish one — not because they lack the data, but because their paperwork forbids it. A rival with a bigger content team cannot outspend you into this; they have to go back and fix a retention schedule written four years ago.
The acquisition pattern is specific. Links come from people who need the number to answer a question they are asked repeatedly: forum moderators, club and association pages, trade bodies, procurement guides and journalists on deadline. The anchor is usually the literal phrase people search for rather than a commercial term. They renew with the calendar — every update earns a fresh round while the previous editions keep the links they already have — which is a very different acquisition curve from campaign-driven work, and it shows up as a slow, steady line rather than the spikes discussed in analyses of link velocity in 2026.
The programme, in order
Write the sentence you want to be cited for before you check whether you can publish it. “Walking boots in this category run half a size small in 58% of exchanges” is a sentence; “we have great returns data” is not. Then work backwards through the four preconditions until you find the one that blocks you, and fix that one only.
Publish the figure on a page that exists to hold it, with the definition, the period and the sample size on the page. If the dataset supports a chart or a small tool, build it — an interactive calculator can earn links for years where a PDF earns them for a fortnight, and the same is true of a well-built scrollytelling explainer when the data has a shape worth walking through.
Then distribute along two tracks. The first is direct: send the URL, not a pitch, to the people who already answer this question manually. The second is journalist-facing, where a genuine first-party number is one of the few things that reliably converts — the HARO-style workflows still work in 2026 if you lead with the dataset rather than an opinion, and the comparison of Connectively, Featured and Qwoted is the practical guide to where to put it. A dated, defensible statistic also makes newsjacking viable for companies that have nothing else to say when a story breaks in their sector.
Two caveats, stated plainly. These links are low volume and mostly from modest domains; they will not move a contested head term alone, and belong in a portfolio alongside the tactics in the link building strategies hub. And a licence question sits underneath all of it: publishing a number openly means machines will ingest it, which is a deliberate trade with its own terms — the ground covered in the AI content licensing and pay-per-crawl playbook. Decide that consciously rather than discovering it later — and per jurisdiction, since the calculus differs across the regimes discussed in international link building.
One line that should not need saying: never invent the number. Fabricated data is the fastest way to lose the only advantage this approach confers, which is that yours is the version that can be checked. If the data is thin, publish the thin version with the n stated and let it grow.
8. The Monday checklist
Seven steps, in order, none of which needs a budget approval to start.
- Write the sentence. One statistic you would want cited, phrased as a complete claim with a period and a denominator. If nobody in the room can write one, you have not identified an outward dataset yet.
- Find the deletion date. Ask engineering, not legal, when the underlying records are purged. This is the most common blocker and the one with the longest lead time.
- Read your own privacy notice. Look for statistical use and external publication in the stated purposes. If they are absent, that is a one-paragraph amendment that should go in this week.
- Check the cookie basis. If the behavioural data you want to publish is being collected under the PECR statistical exception, it cannot be shared. Route the publishable analysis through a different lawful basis, or use order data instead of session data.
- Search your contracts for “aggregated”. Customer DPAs, marketplace terms, platform agreements. Find out whether you hold the data as a controller or on somebody else’s behalf, and whether external publication is permitted.
- Set a minimum cell size and write it down. Fifty is a defensible starting point for most commercial datasets. Record the identifiability assessment: the ICO expects the reasoning to exist, not just the output.
- Publish one page, then wait a quarter. One fact, one URL, one dated definition. Send it to five people who answer that question for a living, and measure referring domains and unlinked mentions of the exact phrase — the measurement approach set out in the 2026 playbook for AI search visibility.
The failure threshold
If, two quarters after publishing, the pages are demonstrably fetchable, the figures are genuinely unavailable elsewhere, and no referring domain and no unlinked mention has appeared, stop. The constraint is demand, not permission: nobody in your market needs to verify that class of fact, and further publication will not change it. Move the budget to earned coverage and conventional acquisition — the fundamentals in the complete guide to link building and the tooling questions covered in the best link building tools round-up. Keep the permission work anyway: the retention and contract fixes cost nothing to maintain.
The broader point survives either outcome. Every privacy control your company owns was built to stop data leaving. In a market where an intermediary decides what your customer sees, the scarce capability is the opposite one: a deliberate, lawful, documented ability to let a fact out under your own name. That capability is not bought from a vendor and it does not show up in the link building statistics for 2026. It is written in a privacy notice, a retention schedule and a contract clause — and, unusually for anything in this field, it is entirely within your control.
