Proprietary Data Moat

Proprietary Data as a Moat Against AI-Generated Competitors

TL;DR

Information gain is measured against the pages already ranking beside you. A moat is measured against a firm that has not entered yet. Optimising the first tells you almost nothing about the second, and the two can move in opposite directions.

What your data cost you is uninformative. The only number that decides durability is the invoice a rival would receive — and on that invoice only two lines cannot be settled in cash: elapsed time, and consent that was already given.

Perishable data defends better than durable data. A rival has to earn back a fixed cost against a stream that is decaying while they build it, which is why the annual flagship study is the weakest shape you can own.

Two instruments: THE RIVAL’S INVOICE, which prices your position from the other side, and THE STANDING-START GAP, which measures how much of your published work a competent rival could restate within one reporting period.

For link building: these positions end by aggregation, not by theft. Sort prospects by aggregation threat rather than authority, and buy reliance rather than reach.

The field found a real effect and drew the wrong conclusion from it

The effect is real and it is now measured. On-Page.ai scored 150 pages holding top-three Google positions across 50 keywords in ten verticals, comparing each page against the cohort ranking beside it by meaning rather than wording. The median page scored 52 out of 100. Roughly one in four scored below 40 — adding almost nothing beyond what already ranked alongside it — while one in five cleared 70. Kevin Indig’s cut of the same dataset is the part everyone quoted: pages carrying at most one unique figure averaged 40.2, and pages carrying fifteen or more averaged 62.1, with the score climbing at every step in between.

So unique numbers pay. That is not in dispute here, and the statistics on what earns citations point the same way. What is in dispute is the sentence the industry attached to the finding. Search Engine Land’s January 2026 panel of search leaders put it plainly: brands are building entity moats by strategically naming their data, and if a competitor cannot replicate your data they are forced to cite your name. Agency guidance through mid-2026 goes further, describing a content moat as a barrier of proprietary assets that is structurally impossible for competitors to replicate.

The problem is the referent. An information gain score compares your page to what exists today. It is a measure of difference from the current cohort. A moat is a claim about what will exist tomorrow — specifically, about what a rival will and will not build. These are not two views of the same property. A number that is unique today precisely because nobody has bothered to produce it scores at the top of the first measure and at the bottom of the second.

What is a moat, exactly?

What is a moat? A moat is a barrier to entry: something that stops a competent, motivated rival from earning the returns you are earning. Bruce Greenwald’s formulation is the useful one — competitive advantage is not doing something well, it is doing something a rival cannot do. That makes a moat a property of a situation rather than a property of an asset. You cannot look at a dataset and say whether it is a moat, any more than you can look at a wall and say whether it is a defence. It depends entirely on who is outside it and what they want.

Proprietary is a statement about access: this is mine, you cannot have it. A moat is a statement about arithmetic, and the arithmetic belongs to somebody else. Two firms publish a benchmark in the same week. One derives it from four years of meter readings on equipment it maintains under contract. The other derives it from a 900-person panel survey fielded in four days. Both score well on originality, both get picked up, and the difference between them is invisible on the page. Only one of them owns something a rival cannot buy on Tuesday.

Your collection cost is uninformative; the rival’s invoice is the whole question

The instinct is to price defensibility by effort — this was hard for us, so it must be hard for them. The error runs in a systematic direction: the data that is cheapest for you to hold is often the most expensive for anyone else to copy, and the data that is most expensive for you to buy is the cheapest for anyone else to buy.

Indig has half of this right — the most defensible numbers are a by-product of the business itself rather than something assembled to feed a content calendar. The reason matters more than the rule. By-product data is not more authentic. It is attached to a position, and the position carries the price. Nobody can produce your maintenance telemetry without maintaining your installed base, and the installed base is what would appear on the invoice.

The commissioned survey is the clean counter-case, and it is worth being blunt about it because it remains the format most agencies sell as original research. UK panel providers price responses from around £2 each, rising past £10 for tight targeting, and a bespoke consumer poll fields in three to four days. The replication cost of your flagship survey is published on a vendor’s website and the lead time is shorter than your editorial calendar. The panel company owns the position; you rented it, and they will rent it again to whoever asks next. Industry reporting in 2026 puts fraudulent responses — bots, professional respondents, click-through incentive farming — at around 31% of raw survey data, so the thing you rented is degrading as well as shared.

The instrument below reprices your position from the outside. It is not an audit of what you own. It is the itemised bill a competent rival would be handed if they decided tomorrow to take your position away from you.

THE RIVAL’S INVOICE

Line itemWhat the rival is actually buyingPriced inEffect on your position
InstrumentationMeters, sensors, logging, the system of record that turns activity into rowsCapex, falling every yearNone. Cheapest line on the bill and getting cheaper
Labour and analysisCollection, cleaning, modelling, writing upDay rates, now heavily automatedNone. This is the line AI actually deleted
Purchased sampleResponses from a vendor’s panel poolAbout £2–£10 per response, 3–4 day lead timeNone. Your survey has a published market price
Access and standingThe right to observe a population you do not ownAcquisition, partnership, or years of tradingPartial. Money buys it slowly, and it can be revoked
Consent and permissionLawful basis to hold and publish observations already takenNot obtainable retrospectivelyReal. Cannot be bought after the fact at any price
Elapsed timeReporting periods of history that had to be lived throughNot for saleReal. The only line no budget compresses

Two lines are green, and both of them are historical. Consent had to be given at the time it was given. Time had to pass. Everything a rival cannot buy is something that had to happen in the past — which means a position built entirely out of present-tense advantages is not a moat at all. It is a head start with good presentation, and the whole link building playbook is currently being pointed at head starts.

Key takeaway

Price your data position by what a rival would have to spend, not by what you spent. Four of the six lines on that invoice are settled in cash and three of them are falling in price annually. If your advantage sits on those four lines, you are holding a lead measured in months.

Durability is a race between two clocks

Once the invoice is priced, durability becomes a comparison between two clocks that most teams never put side by side. The first is the entrant’s payback clock: how long it takes a rival to earn back the cost of building an equivalent position. The second is the shelf-life clock: how long a single period’s numbers stay relevant to the decision they inform.

Entry is rational when payback lands inside the useful life of what is being bought. Deterrence happens when it does not. That comparison produces a result that runs directly against instinct.

Perishable data defends better than durable data

If your numbers go stale in eighteen months, a rival must pay the full fixed cost to acquire a stream that is already decaying while they build it. You, with the meters installed and the process running, pay only the marginal cost of the next period. That asymmetry — a sunk fixed cost on one side and a low marginal cost on the other — is the oldest entry deterrent in industrial economics, and it is as available to a twenty-person firm as to a large one. Fast decay plus a standing apparatus is a stronger defence than a slow-decaying dataset held by nobody in particular.

The inverse is the finding worth taking away. A durable, one-off dataset is the worst shape you can own. The entrant pays once and holds it forever, which is exactly the shape of the discipline’s favourite format. An annual flagship study is a fixed cost paid every twelve months to defend a position for about six weeks, and what it produces is a stock any competitor can match with a single purchase. It is excellent press and a poor asset. A recruitment firm publishing time-to-hire from its own applicant system every month holds something a rival cannot start; the same firm publishing an annual salary survey holds something a rival can order.

Key takeaway

The defensible unit is not the study, it is the meter. Cadence beats magnitude: a small quarterly series with a fixed definition outperforms a sixty-page annual report, because the report is a stock with a market price and the series is a flow with a start date somebody else missed.

The diagnostic problem: not worth copying looks exactly like impossible to copy

From inside a business, two completely different situations present identically. In the first, nobody has replicated your dataset because they cannot. In the second, nobody has replicated it because there is no reason to. Both look like: we are the only people who have this. Only one is a moat; the other is an unsold product with a storage cost. Separating them requires two measurements — one on the supply side, one on the demand side — and almost every content team runs only the first.

THE STANDING-START GAP

1. Inventory the claims. List every published sentence on your site that rests on data you generated. Work at sentence level, not page level. Most firms find between 20 and 60.

2. Count the periods. For each claim, write down the minimum number of reporting periods of history required to state it. A survey finding is one. A trend is at least three. A seasonal relationship is at least two full cycles.

3. Mark the access class. For each claim, ask whether a rival could observe the same population without buying something unbuyable. A panel is purchasable. Your own installed base, claim file or transaction log is not.

4. Compute the gap. The Standing-Start Gap is the share of your claims that a well-funded rival could not restate within one reporting period, starting today with no history. Below 20%: exposed — you have information gain and no moat. 20–50%: mixed. Above 50%: defended.

5. Run the demand check. For each defended claim, name the third party whose own document would have to change if your number changed. If you cannot name one, you may be holding something nobody replicates because nobody needs it — and that is the failure mode this test exists to catch.

Then the arithmetic. Price the four purchasable lines of the invoice for your strongest claim using published market rates, divide by what the position is plausibly worth to a rival per year, and compare the resulting payback in years against the shelf life of one period’s data. Payback shorter than shelf life: expect entry. Longer: you are protected by the clock rather than by your cleverness — which is still protection, and worth knowing you depend on.

The fifth step is the one that changes budgets. Supply-side scarcity with no demand behind it is a hobby, and it is remarkably easy to fund for years because it produces the same internal signal as a genuine position: nobody else has this. Naming the downstream document is the cheapest available correction, and it is also, not coincidentally, a prospecting list.

Aggregation is how these positions actually end

Data positions rarely die by theft. They die by aggregation: a body with standing over many operators starts collecting the same field from all of them, and your exclusive becomes one row inside somebody else’s table. What makes it fatal rather than merely competitive is that the aggregator’s version strictly dominates yours — larger sample, wider coverage, and, crucially for citation, visibly disinterested. You are a party with a commercial stake in the number. They are not. Nothing in your writing closes that gap, which is why the usual advice about demonstrating authenticity has no purchase on this particular failure.

UK subsidence shows the shape clearly. The Association of British Insurers publishes a quarterly Property Insurance Tracker; its Q2 2026 release put domestic subsidence payouts at £72m for the quarter with the average claim reaching a record £20,000, against a record £307m across 2025 and £219m in the 2022 surge year. No engineering practice in Britain will out-cite the ABI on claim costs. But the ABI does not publish property-level seasonal movement, because it does not hold it. An aggregator defines the field it collects, and every field it does not collect is still available.

THE AGGREGATOR SCREEN

For each dataset you hold, name the body that would end your position if it began collecting the same field: a trade association, a regulator, an insurer panel, a standards committee, a comparison intermediary, a dominant platform.

That name is simultaneously your largest threat and your best placement target. The cheapest way to survive an aggregator is to become the source of the field they lack and to be named inside their methodology, rather than publishing a competing headline against them.

An aggregator who cites you extends your position. An aggregator who collects your field ends it. Which of the two happens is usually decided by whether you approached them first.

The second ending is legislative

The Data (Use and Access) Act 2025 received Royal Assent on 19 June 2025 and gives government the power to compel firms in any sector to share both customer data and business data — pricing and product information included — under smart data schemes. The Smart Data 2035 Strategy, presented to Parliament in March 2026, targets five or more active schemes by 2030 and twenty or more by 2035, backed by at least £36m over four years, with banking, finance, energy, property, retail, transport, telecoms, digital markets and agrifood all named. In parallel, manufacturers selling connected products into the EU have until 12 September 2026 to build data access into those products by design under the EU Data Act.

If your data comes from a connected product or a customer relationship in one of those sectors, your exclusivity is now a policy variable with a published roadmap and a date attached. That is not an argument for skipping the work — it is an argument for preferring positions built on observation you perform over data your customers generate, because the second kind is being unbundled by statute. Teams already tracking European content regulation should be reading the smart data roadmap with the same attention.

Key takeaway

Model the aggregator before you model the competitor. In most UK sectors the party capable of ending a data position is not a rival firm at all — it is a trade body, a regulator or a statutory sharing scheme, and all three are approachable in a way a competitor is not.

Productising: turning a sunk cost into something a rival must match every period

The 2026 consensus on productisation stops at naming — own a metric, call it the [Brand] Index, and models cannot ignore it. Naming genuinely helps, and it is also cheap, universal within two years, and entirely copyable. Naming is a distribution device. It is not a barrier. Three things make a named series hard to displace, in ascending order of durability.

1. A frozen definition

Should you ever change the methodology? Almost never, and this is the least intuitive rule in the whole discipline. Every change to a definition resets the back-run to zero, because a series with a discontinuity cannot support a trend claim across the break — and the trend claim is the only part of your output a new entrant structurally cannot make. The most expensive thing you can do to a proprietary dataset is improve it. If a change is unavoidable, run both definitions in parallel for a full cycle and publish the conversion.

2. A published cadence

A date is a commitment. A quarterly release with a stated next-publication date tells a downstream user that the number will still exist when they need it again, which is the precondition for them building anything on top of it. Ad-hoc publication, however good, cannot be depended on, and dependence is the entire game. Machine-readable delivery matters here too: an API or structured feed makes your series usable inside somebody else’s system rather than merely readable to a browsing agent on your own page.

3. Adoption by someone who will have to rewrite

The moment a specifier, an underwriter, a guidance note or a procurement template contains your figure, an equally good rival number stops being sufficient. It now has to be better by enough to justify rewriting a document, retraining the people who use it and defending the change to whoever signed the last version. That is switching cost, and it is the only one of the three that a competitor cannot acquire by copying you. On the publication question itself: release the finding at full confidence and hold the row-level file, or license it on terms. The reason is in the objection below.

What this changes about how you earn links

Four changes, none of which are about domain authority and none of which show up in a standard prospecting workflow. They sit on top of the fundamentals of earning links, not beside them.

Sort prospects by aggregation threat, not by authority score

The bodies most capable of ending your data position are the ones most worth placing with, and they are systematically underweighted by every tool in the link building toolset because they score poorly on traffic and topical relevance. A standards committee, an insurer’s technical panel or a trade association’s guidance note is usually a weak link on paper and a strong position in practice. This is closer in spirit to sponsorship and institutional placement than to editorial outreach, and it is measured in adoption rather than in entity authority movement.

Buy reliance, not reach

A link from a page that repeats your number is counted as a citation and little else. A link from a document that depends on your number is a switching cost. The second is worth several of the first and it is bought differently: you have to make the figure usable inside the other party’s process — a threshold, a band, a formula, a downloadable table — rather than merely interesting to their readers. A pest control firm whose seasonal activity index becomes the trigger table inside a facilities management provider’s inspection schedule has bought something that a guest placement or a listicle inclusion cannot deliver at any volume, however well those channels perform on recommendation surfaces.

Old links appreciate, for a reason nobody cites

You can back-date your own archive. You cannot back-date somebody else’s page. A third-party document from 2022 that cites your series is the only externally dated proof that the back-run existed — and the back-run is precisely the part of the asset a rival cannot buy. This inverts the usual case for aged links: the value is not accumulated authority but timestamping. It also means that recovering lost historical citations is an asset-protection exercise rather than a traffic one, which changes how you should budget citation recovery work and how you treat edits into existing published pages that carry old dates.

If you hold no data, buy a field in someone else’s series

Most firms have no proprietary dataset and no realistic path to one. The invoice still tells you what to shop for. Supplying a single column to an existing benchmark — one field, consistently, with your name in the methodology — is dramatically cheaper than building a series and delivers the same two green lines: your contribution accrues history, and your consent to be named is already recorded. For firms working across multiple national markets, the same logic applies per market, since aggregators are almost always national.

Where this argument is weakest

The strongest objection is not that the invoice is wrong. It is that the invoice is falling. Three of the six lines — instrumentation, labour and purchased sample — get cheaper every year, and the green lines are now under direct attack. A model with public soil maps, rainfall records and claims aggregates can estimate a seasonal movement curve without monitoring a single property. If synthetic reconstruction can approximate a back-run, then elapsed time stops being unbuyable and the whole framework loses its last defended line.

That objection is correct in direction and should be conceded without hedging. It bounds in four places.

  • Imputation reproduces the centre and fails in the tail. Every decision anyone actually pays for is a tail decision — not what clay does on average, but what this house does this winter. A reconstructed curve is a prior, and priors are exactly what a party consults you to override.
  • The binding constraint is contractual, not epistemic. Once your figure sits inside somebody’s process, a substitute has to be better by enough to justify a rewrite. Reconstruction produces a rival number, not a reason to switch, and switching is where the money is.
  • Reconstruction needs a training signal, which means it needs your rows. Publishing the full row-level file at high resolution is how a replication gets financed. Publish the finding; hold or license the file. This is the one place where how your material enters training corpora is a commercial decision rather than a philosophical one.
  • Honest bound: this is a claim about the next two to three years. The cost curve is moving in the entrant’s favour on most lines and there is no reason to assume it stops. A framework whose green rows are shrinking should be used to pick which position to build now, not to justify a ten-year plan.

A second objection deserves a straight answer: most readers hold nothing proprietary and never will. Conceded. The cheapest green line available to a small firm is not collection at all — it is starting a clock. Publish one small measurement this quarter under a definition you intend to keep, and get it into a third party’s document before the year ends. A moat of this kind is built by the calendar. The only requirement is not stopping.

Worked example: nine winters that nobody else has

Ridgemont Structural is a Chelmsford structural engineering practice, £6.2M in annual fees, 41 staff, sitting on three insurer panels for domestic subsidence across the Essex and Hertfordshire clay belt. The figures are illustrative; the mechanics and published anchors are not.

The standard play, 2025. Ridgemont commissioned a 1,000-respondent homeowner survey — subsidence worry, awareness of cover, willingness to remove a tree — at £24k all in, published as an annual report. It earned 31 referring domains and a fortnight of trade coverage. In April 2026 a national underpinning contractor published a near-identical survey with a larger sample, and Ridgemont’s report stopped being cited inside a quarter. Replication cost, at published panel rates: roughly £11k and eleven days.

The audit, February 2026. Running the Standing-Start Gap across the site found 38 claims resting on Ridgemont’s own data, of which 34 could be restated by a funded rival within one reporting period. Standing-Start Gap: 11%. The same audit surfaced what had never been published — quarterly crack and level monitoring on 3,140 properties since 2017, held only because insurers require monitoring across a full seasonal cycle before agreeing remediation. The Financial Ombudsman Service treats twelve months of monitoring as reasonable precisely because it covers all four seasons, and BRE Digest 251 supplies the classification convention. The elapsed-time line on the invoice was, in this trade, written into the claims process itself.

The build, March to July 2026. One thing was published: the Clay Belt Movement Index — seasonal movement amplitude banded by soil plasticity and by distance to mature trees as a ratio of tree height, 2017 to 2026, updated quarterly, definition frozen and printed alongside every release. The headline finding was that on high-plasticity clay, properties inside 0.6 times mature tree height showed a median seasonal amplitude of 4.1mm against 1.3mm beyond 1.2 times, and that 2025 amplitude exceeded the 2018 surge year at 61% of monitored properties. Nine years of history are required to state that sentence. Nobody starting in 2026 can state it in 2027, or in 2030.

The aggregator screen. Two bodies could have ended the position: the ABI, which holds claim cost, and an arboricultural association, which holds tree management practice. Neither collects property-level amplitude. Ridgemont supplied the index as a named input to the arboricultural body’s guidance rather than publishing a competing note against it. By August 2026 the index was cited in two insurers’ internal contractor guidance and specified in one loss adjuster’s instruction template as the reference for pre-works monitoring bands. Twelve external documents cited it; four would require rewriting if the definition changed.

Four things that went wrong

  • The greenest line turned out to be contractual. A June 2026 panel renewal asserted the insurer’s ownership of monitoring data generated under instruction, and Ridgemont lost a fifth of its properties from that quarter onward. Standing is granted, and it is withdrawn by a clause rather than by a competitor. Publication rights now get negotiated at renewal, in advance, as a condition of panel membership.
  • The methodology freeze cost real money. A migration from tell-tale gauges to digital displacement sensors, begun in 2024, produced readings that are not directly comparable. Preserving the back-run required running both instruments in parallel on 300 properties for four quarters — about £38k of pure duplication — and publishing the conversion. Improving the instrument nearly destroyed the asset.
  • The moat is a subscription, not an asset. Amplitude data stays decision-relevant for roughly eighteen months, which is what deters entry and also means the £70k a year of collection is not optional. One bad trading year silts the moat in a single period, and the framework offers nothing against that beyond saying it out loud before anyone commits.
  • Defensibility and citability are only loosely related. Through 2026 the most-quoted Ridgemont sentence in AI answers remained a one-period homeowner-behaviour figure from the survey — the thing any rival could buy. The defended claims earned fewer mentions and all four of the documents. This framework promises durability, not popularity, and anyone measured on mention volume will experience the correct strategy as a downgrade for at least two quarters.

The Monday checklist

  • List every published claim that rests on your own data, and mark each with the minimum periods of history required to state it. Compute your Standing-Start Gap before you commission anything new.
  • Price the four purchasable lines of the rival’s invoice for your strongest claim, using vendor rates you can look up this afternoon. If the total is under a quarter’s content budget, stop calling it a moat.
  • Name the aggregator for each dataset you hold — trade body, regulator, insurer panel, intermediary — and put that name at the top of your prospecting list rather than at the bottom.
  • Freeze one definition and publish a next-release date beside it. If a change is unavoidable, run both definitions in parallel for a full cycle and publish the conversion.
  • Identify one document belonging to someone else that could contain your figure as a threshold, band or formula, and rebuild the output to fit inside it.
  • Check whether your sector appears in the smart data roadmap. If it does, shift weight toward observations you perform rather than data your customers generate.
  • If you hold no dataset at all, pick one measurement you can repeat every quarter for three years, publish the first one this month, and place it somewhere dated that you do not control. Everything defensible on that invoice starts as something that had to begin earlier than now.

The sentence worth keeping is the one that survives the whole framework: a moat is not in your data — it is in other people’s documents. Data is what makes the entry possible. Reliance is what makes leaving expensive. For a link building specialist, that is a familiar shape in unfamiliar clothing, and it is the reason this problem lands on your desk rather than the analytics team’s.

Leave a Reply

Your email address will not be published. Required fields are marked *

First-Hand Experience Signals Previous post First-Hand Experience Signals: The E-E-A-T Input AI Can’t Manufacture
AI-Generated Competitor Content Next post Detecting and Out-Competing AI-Generated Competitor Content