Provenance-Backed Original Research

Provenance-Backed Original Research: Data That Can’t Be Faked

TL;DR

A provenance chain begins at the moment of capture. Every choice that decides whether a commercial statistic is true — which units you measured, from what vantage, under whose definition — is made before the first byte exists, and no signature reaches any of them.

A fabricated dataset can carry a flawless chain of custody, and if somebody fabricated it, it will. The manifest is the cheap part.

The right to check almost never gets exercised. Where journals made data-availability statements mandatory, 93% of authors who had promised in print to share their data did not, and 6.8% delivered — the same rate as authors who promised nothing.

Re-derivability and credit pull against each other. A number a stranger can recompute gets restated without you; a number only you hold must be credited to you and cannot be checked.

So buy a completed check rather than an offer of one: a population you do not control, a second vantage that sees what your instrument cannot, and one named third party who has already re-run your figures and published whatever they got.

Then name the measurement. You cannot sign a sentence, but you can name a number, and the name travels into every restatement.

1. The promise, stated at its strongest

The proposition goes like this. Sign your dataset. Publish the manifest. Deposit the file somewhere permanent with an identifier. Your research then becomes the source of record in a web where most numbers are unverifiable, and the engines and journalists who need a figure will reach for the one that can be proved. Unfakeable, therefore trusted, therefore cited, therefore linked.

What is provenance-backed original research?

Original research published together with cryptographic or metadata evidence of where the data came from and what has happened to it since — signed manifests, published hashes, lineage records, a deposited copy with a permanent identifier. The category is new enough that most of what is written about it treats provenance as a trust signal by analogy with a certificate, without asking which part of the research it certifies.

The tooling behind that pitch is real and current. Croissant, the metadata format for machine-learning datasets, reached version 1.1 on 12 February 2026, adding machine-actionable provenance built on the W3C PROV-O model for lineage and ODRL for usage conditions. Adoption is not theoretical: over 700,000 datasets carry Croissant metadata, Hugging Face, Kaggle, OpenML and Harvard Dataverse have implemented it, Google Dataset Search offers a Croissant-only filter, and NeurIPS now requires the metadata in dataset-track submissions.

What does a provenance chain actually prove?

That a file has not changed since a named party signed it at a stated time. It is a statement about custody. It is not a statement about measurement, and the two are so easily confused that an entire content-marketing thesis has been built on the confusion.

Custody begins at capture. Write out the reasons anyone has ever had to distrust a commercial statistic — the sample was drawn from the client’s own customer list, the question was worded to produce the answer, the units that would have spoiled the result were excluded for a defensible-sounding reason, the denominator was chosen after the numerator was known — and notice that every single one of them is settled before there is a byte to sign. The failure mode that a chain of custody defends against, somebody quietly altering your published figures after release, is close to unknown in this field. The failure modes that are everywhere are choices, and they all sit upstream of the seal.

There is a second-order problem too. A competent fabricator produces the best provenance in the room, because signing is cheap, it is the only part a machine can check without effort, and it is entirely compatible with an invented dataset. If you build a regime that rewards manifests, you raise the floor on presentation and leave the floor on truth exactly where it was. That is worth holding in mind before spending a data-led campaign budget on cryptography.

2. Two numbers in the same box

The clearest live demonstration in British commerce is bolted to the side of a motorway services car park.

The Public Charge Point Regulations 2023 came into force on 24 November 2023. Since November 2024, every rapid charging network in Great Britain — anything at 50kW and above — must average 99% reliability across the estate, measured annually. Reliability is not self-described in prose: it is computed from EVSE status objects in the Open Charge Point Interface, the mandated data standard, where available, charging and reserved count as up while blocked and unknown count against you. Real-time status must be published free of charge in machine-readable form. Operators must publish their compliance position on their own websites. The first annual reports landed in January 2026, enforcement sits with the Office for Product Safety and Standards, and the penalty for missing the reliability standard is £10,000 at network level.

That is a stronger provenance position than any content-marketing dataset will ever hold: instrumented capture, a protocol-defined schema, a statutory reporting duty, mandatory open publication, and a regulator with a fine attached.

Now look at the other number the same box produces. To bill a driver by the kilowatt-hour, the energy meter inside that charger must be certified under the Measuring Instruments Regulations 2016 — independently tested by a notified body, sealed against tampering, and marked CE M with the body’s identification number. An uncertified internal meter gives an operator a reading it can look at. A certified one gives a reading that can be enforced against a paying customer.

So the charger holds two numbers: one that an independent body tested to a defined accuracy class, and one that the operator defined for itself. Only the second one ever appears in a press release.

ChargerHelp analysed more than 100,000 charging sessions across 2,400 chargers and found that networks reporting 98.7–99.9% uptime delivered a first-time charge success rate of 71%, with more than a third of failures occurring on chargers that appeared to be working. Monta’s 2025 survey of over 200 UK charge point operator decision-makers found 3.9% believed they currently met the 99% threshold.

Nothing in that gap is fraud, and nothing in it would be caught by a manifest. Every byte in the pipeline is exactly what the charger reported. The figure is wrong about the thing drivers care about because the instrument reports its own status, and the failure that matters most consists of a machine not working while believing that it works.

Key takeaway

The best-attested availability figure in British retail energy sits roughly 28 points away from user-observed reality, with an unbroken chain of custody the whole way. Integrity was never the binding constraint on whether a number was worth believing.

3. The three choices no signature reaches

Three decisions determine what a statistic means. All three are made by a human before capture, none of them leaves a cryptographic trace, and together they account for almost every published figure that later turns out to be indefensible.

The population

Which units enter the frame, and which are quietly outside it. The charging regulations exempt periods when a physical barrier blocks access to a charger, and carry a reasonable-excuse provision on top. Neither is dishonest; both mean the frame was drawn by the party being measured. For example, a software firm publishing an average onboarding time from its own instance data has, without any intent to mislead, excluded every customer who gave up during onboarding — which is to say it has excluded the finding.

The vantage

Where the instrument sits, and therefore what it structurally cannot see. A charger cannot log the session that never started. An analytics tag cannot see the visit that ended at a paywall. A scraper does not observe prices; it observes a page, which makes the resulting dataset a report of a report. Choosing a vantage is choosing a blind spot, and the blind spot never appears in the output as a gap — it appears as a clean number.

The definition

What counts as an event, and what sits underneath as the denominator. Counting reserved as available is a definitional choice with a double-digit consequence. In survey work the equivalent is whether considering includes might, eventually. A denominator is a definition wearing arithmetic.

Set the chain out in full and the split is obvious. The head decides the number. The tail is where all the tooling lives.

StageWhat it decidesWhat provenance can attestWhat an outsider needs to re-derive it
PopulationWhich units are in the frame, and which are excludedNothingThe frame, every exclusion, and a route to the units you left out
VantageWhere the instrument sits, so what it cannot observeNothingA second instrument that sees the same event differently
DefinitionWhat counts as an event, and the denominatorNothingThe written rule, tight enough that two strangers compute alike
CaptureThat a reading was taken, when, and by what deviceTime, device, integrity of the reading as takenThe raw extract rather than the summary
ProcessingThe steps from raw reading to published statisticLineage: what ran, in what order, on what inputThe code, or the arithmetic written out in prose
PublicationThat the file you can download is the file that was signedFull integrity and non-repudiationNothing further. This is the part already solved

The Measurement Chain. The three red stages decide whether the figure is true. Not one of them has a field in any provenance format.

4. Almost nobody exercises the right to check

The second assumption inside the pitch is that making data available makes it checked. This one has been measured, at scale, in the setting most favourable to it.

Gabelica, Bojcic and Puljak examined 3,556 articles published in a single month across 333 open-access journals operating a mandatory data-availability statement. Nearly all — 3,416 — carried one, and the most common category, at 42%, said the data were available on reasonable request. The researchers then asked. Of the 1,792 manuscripts whose authors had stated a willingness to share, 1,670 either did not reply or refused: 93%. Fourteen per cent responded at all. Six point eight per cent supplied the data. The authors’ own conclusion was that compliance among those who promised was no better than among those who promised nothing.

A 2026 study of 157 meta-analyses in highly ranked sports journals found the same shape from the other direction: 34% carried a data-availability statement, 33% shared data, 11% shared code, and the presence of a statement was not associated with greater sharing.

Nor does the verification happen in the population most professionally obliged to do it. The GhostCite work published in 2026 surveyed researchers and found 87% claiming they always verify AI-generated citations while 42% paste reference entries in without checking and 77% of reviewers do not check references thoroughly. Retraction Watch reported in May 2026 that roughly one in 277 PubMed-indexed papers this year contains fabricated references, and the commentary around it made the mechanism plain: the infrastructure propagates, it does not audit.

Machines are no better yet. An Oumi analysis of Gemini-3-powered AI Overviews found 91% contained the correct answer while only 39% were both correct and fully supported by the sources cited — which is a polite way of saying the citation was attached rather than checked.

Does publishing your dataset make your research more citable?

Only if somebody opens it. Availability is an option, and an option is worth its exercise rate. In the best-documented setting available to us, that rate is about seven per cent. An offer to share is not a signal; it is a promise, and promises are priced at their observed keeping rate.

5. The carve-out that gives the game away

If you commission survey research in Britain you can buy a credibility badge with it: a member of the British Polling Council, abiding by its rules. Those rules are genuinely demanding. Within two working days of publication a member must place on its own website a full description of the sampling procedures, computer tables showing the exact questions in the order they were asked with all response codes, weighted and unweighted bases for every demographic published, and a description of the weighting applied.

Then read the carve-outs. Polls conducted to support the public relations activity of a commercial client need not be published within two working days; they revert to publication on request. And the disclosure rules do not apply at all where the survey organisation had no responsibility for the design of the survey or the analysis, and all weighting of the data followed client instruction.

Put those together and the badge on a PR survey attests to fieldwork. The three things that decide the number — who drew the frame, who wrote the questions, who chose the weighting — can all belong to the client, and the duty to publish proactively does not attach. This is not an accusation aimed at pollsters. It is the trade body’s own rule, written by people who understand exactly where in the chain the discretion lives, and it draws the line in precisely the same place the mechanics do.

The operational version is three questions, in writing, to any research supplier before the invoice: who wrote the questionnaire and who signed off the sampling frame; whether full tables will sit at a public URL on the day of release rather than on request; and whether the supplier will take a journalist’s methodology question directly instead of routing it back through you. A supplier who says yes to all three is selling you something a digital PR campaign can stand on. A supplier who says no to all three is selling fieldwork with a logo.

Key takeaway

The credibility badge on commissioned research usually certifies the part nobody disputes. Establish what the badge covers, then buy the part it does not.

6. The trade between being checkable and being credited

Here is the part the playbook never prices, and it is the reason sensible teams keep making the wrong call.

A figure a stranger can recompute is a figure a stranger can restate flatly. It needs no attribution, because in an important sense it is not yours — you reached it first. Numbers derived from public registers, freedom-of-information returns and statutory filings are simultaneously the most credible things you can publish and the most easily orphaned, because the second writer can go to the same source and say the same thing without mentioning you.

A figure only you can produce forces attribution. Nobody can restate it without you, because nobody else can reach it. But the property that compels the credit is the property that makes it uncheckable, and an uncheckable number arrives wrapped in a hedge — according to — which in written media is a citation and in a generated answer is the first thing stripped out.

That asymmetry deserves a moment, because it runs in opposite directions on the two surfaces that matter. In text, a hedge is good news: the writer who is not certain of your number has to say whose number it is, and the hedge carries your name and often a link. In a generated answer the hedge is overhead, and attribution is the first thing compression removes, so the proprietary figure loses precisely the credit it was supposed to compel. If you are auditing where you appear and your data-led work is invisible, the reason is often this rather than anything about the page.

So the two goals pull apart. Credibility argues for re-derivable data. Credit argues for proprietary data. Most teams resolve this by accident, publish something proprietary, and then measure the wrong thing, counting placements when the interesting number is what happened to the figure afterwards.

The Restatement Ledger

For each headline figure, search the exact number and its distinctive phrasing across the 90 days after release, and classify every instance you find: linked, named without a link, or orphaned. Then two rates.

Restatement Yield = restatements ÷ placements you secured. Below 1.0 the figure is not travelling and you bought coverage, not a reference. Above 3.0 it is travelling well, which is only good news if retention holds up.

Attribution Retention = (linked + named) ÷ restatements. Below 40%, the number has left home and is now general knowledge. Between 40% and 70% is normal for an unnamed statistic. Above 70% means your name is attached to the measurement itself and not merely to the announcement.

No rank tracker will run that ledger for you, and it is worth checking what your existing toolset can export before doing it by hand: the work is exact-phrase and exact-number matching, which is scriptable but not standard. The resolution itself is nomenclature, not cryptography. A number can be copied; a named measurement carries its owner into every restatement, which is what an index is for and why a house-price index outlives by decades the press release that launched it. Publish the definition as its own standing page — numerator, denominator, exclusions, the classes of failure it counts — separately from the report, and give the measurement a name containing yours. Then keep the definition still while the number moves. You cannot sign a sentence, but you can name a number, and a name survives the paraphrase that strips every other marker of origin. That page also becomes the most linkable technical asset in the programme, because it is the thing a writer needs when they have to explain what your figure means.

7. Buy a completed check, not the right to one

If custody is not the constraint and availability is not exercised, what actually makes a figure trustworthy to somebody who has no reason to trust you? Exposure to contradiction. Three purchases deliver it, cheapest first.

A population you do not control

Freedom-of-information returns, statutory registers, published compliance filings, planning records, market prices. Someone else holds the records, and someone else can pull them again tomorrow. You are not asking to be trusted; you are pointing. This is also the cheapest research any in-house team can run, and its weakness is exactly the orphaning problem above, which is why it needs a name attached from day one.

A second vantage

If your instrument cannot see failure, buy one that can: a small instrumented panel, a mystery-shopping sample, a customer’s own telemetry, a subset re-measured by hand. A second vantage converts a self-report into an observation. And the gap between the two vantages is usually the most publishable thing the business will produce all year, because it is a finding about the world rather than a summary of your own records.

A completed, published check by a named party

This is the only form of verification with a demand side, because somebody has already done it. Commission a named third party — a university group, a chartered auditor, a trade body’s technical committee — to re-run one period from your raw extract and publish whatever they get, with no approval step and no right of reply beyond a factual correction. Price it as a fee for an opinion, not a fee for a result. If the contract lets you bury the answer, you have bought nothing, and everyone downstream will assume you exercised the option.

Depositing the extract in a repository with a persistent identifier is still worth doing, because it puts a copy beyond your own future revisions and it is what makes the third party’s work possible. Just cost it honestly: it is insurance and infrastructure, not a credibility programme. The same goes for machine-readable publication of the underlying tables — necessary, cheap, and not persuasive on its own. Put the extract somewhere a technical audience will actually pull it apart rather than only somewhere tidy: a dataset posted to a community that argues with numbers gets checked, and a data-visualisation build gets admired.

The Second-Path Test

1. Could somebody outside your business reach the same population without your permission?

2. Could they observe the failures your instrument cannot see?

3. Is the definition written tightly enough that two strangers would compute the same figure from the same extract?

4. Does anyone who would lose money if the figure were wrong already have everything they need to say so?

A number is re-checkable when someone with a motive to contradict you holds everything required to do it. Four noes means you have a claim, not a finding — and a claim is still publishable, provided you stop calling it evidence.

8. Where this argument is weakest

The strongest objection is technical, and it is genuinely strong: hardware attestation moves the chain below the human. A sealed and certified meter, a camera that signs at the sensor, a tamper-evident log — these attest to a measurement rather than merely to a file. The certified DC meter inside a rapid charger is exactly this. So is capture signing shipping in cameras from several manufacturers and now in certified handsets. It is the only development anywhere near the head of the chain, and for a physical quantity it works.

Four bounds, in descending order of comfort.

  • It attests the reading, never the placement. Where the meter goes, which sites are metered, and which readings are retained all sit outside the seal.
  • Almost no marketing dataset is a physical quantity. It is a survey, a scrape, a query against your own database, or a model output. A scrape has no sensor at all; it is a report of a report, and the thing it reports on can change under it.
  • Certification is scoped to the quantity certified. The same sealed box also emits an availability figure its owner defined, and that is the figure that ends up in the headline.
  • It raises the cost of the expensive fake and leaves the cheap one untouched. Nobody fabricates meter readings. They choose the sample.

The second objection is that verification is about to become free because it will be machine-run: Croissant’s machine-actionable provenance, PROV-O lineage graphs, a submission requirement at a major conference, governance tags designed to be read by agents rather than people. Concede the trajectory — at some point a retrieval layer will check a manifest because checking costs nothing.

But a gate can only test what a field asserts, and the vocabulary describes lineage, licensing and structure. There is no field for the sampling frame, no field for the vantage, and no field for who wrote the questionnaire. A machine gate will therefore filter the altered and pass the fabricated, which is close to the opposite of the sorting the market needs. And the party best equipped to emit flawless, complete, machine-actionable provenance at industrial scale is a synthetic-content operation, for which the metadata is a build artefact rather than a cost. Any regime that treats a clean manifest as evidence of care will find itself sorting in favour of whoever automated first — the same dynamic that made AI content labelling a disclosure exercise rather than a quality one, and the same logic applies to whether a dataset gets ingested as a training source at all.

9. What this looked like in practice

Wraymoor Charge Network operates 1,140 public devices from Derby on turnover of £21.4M, of which 412 are rapid units — enough to bring the whole estate inside the reliability duty. In January 2026 it filed its first statutory reliability report at 99.2% across the rapid network.

Marketing saw an asset. The Wraymoor Rapid Charging Report cost £34,000, of which £11,000 went on a provenance layer: signed manifests on every published dataset and chart, a published hash of each monthly extract, and a methodology PDF. The headline was that one rapid session in nine ends without a charge — a 9.4% failure rate, computed from the same OCPI status feed as the 99.2%.

Fourteen weeks later: 31 placements, 68 restatements, 12 links. Twenty-eight of the 68 restatements named or linked Wraymoor, an Attribution Retention of 41% against a Restatement Yield of 2.2. Two parties requested the dataset, both competitors. The signing service logged no verifications at all.

The Measurement Chain diagnosis took an afternoon. Both figures came from one instrument. The population excluded three sites under the physical-barrier exemption. The vantage was the charger’s self-reported status, so sessions that failed at handshake never entered the data. The definition counted reserved as available. On the Second-Path Test the answer was no four times.

The rebuild took twelve weeks and £26,000. First, a second vantage: app-side session telemetry plus a 90-day instrumented panel of 40 vehicles run by a logistics customer, which sees the failure a charger cannot. Second, a standing definition page for the Wraymoor Session Success Rate, listing the numerator, the denominator, four classes of failure and every exclusion. Third, a commissioned re-run — a university transport group took the raw extract and the panel data for one quarter and published its own figure with no sign-off.

The re-run came back at 88.1% session success against Wraymoor’s published 90.6%. The 2.5-point gap traced to a time-zone boundary in how overnight sessions were bucketed. Wraymoor published the correction with the university’s number as the headline.

In the five weeks that followed, the correction earned 23 links against the original report’s 12 in fourteen weeks. Attribution Retention on the corrected figure reached 84%, because there was now a named measurement with a page to point at. Naming across 40 buyer prompts moved from 6 to 27. Two fleet tenders quoted the definition page rather than the report.

Four costs, all real.

  • A county-council framework scored them below a rival whose unaudited 99.4% looked better on paper. The examined number lost to the unexamined one, and that will keep happening.
  • The definition page let an aggregator recompute the same measure across four networks. Wraymoor came third. They had written the ruler that ranked them.
  • The panel costs £4,100 a quarter, and a series is only worth anything if it never stops.
  • The re-run took nine weeks, during which the communications calendar had nothing to announce.

The artefact that ended up being cited was the correction, because it was the only figure in the sequence that had survived somebody trying to break it. That is worth sitting with: the link velocity on a published correction beat the launch it corrected, and the reason is not novelty. A correction is the only document in a research programme that carries evidence of a check having actually happened.

10. What this changes about acquisition

Coverage that repeats your number and coverage that examines it are different goods sold at the same price. The first decays into general knowledge, which is what an Attribution Retention below 40% is measuring. The second becomes the reference, and it is easier to acquire than it sounds, because you are offering a publisher or an institution something to do rather than something to run. A trade title that re-runs one of your figures has produced its own article and will link to the source data as a matter of course; a listicle placement that quotes the figure has produced a sentence.

So add one column to the prospecting sheet, alongside whatever you already track about a domain’s referring profile: does a placement here create a party who has checked something, or a party who has repeated something? The second is worth having. The first is worth paying more for, and almost nobody is bidding against you for it, because the field has spent two years buying mentions rather than examinations. It is the oldest principle in how links are actually earned wearing unfamiliar clothes: give somebody a reason to write, not a line to quote.

The Monday checklist

  • Take your most-quoted figure and write its population, vantage and definition in three sentences. Any of the three that is a choice your own business made is the part to fix first.
  • Run the Restatement Ledger on that figure over the last 90 days. Compute both rates before you commission anything new.
  • Send your research supplier the three questions: who designed the questionnaire and frame, will full tables be at a public URL on release day, and will they field methodology questions directly.
  • Name the second vantage — the instrument that would see what yours cannot — and get a price for 90 days of it.
  • Commission one re-run: one period, one named party, published unedited. Take the money from the research line, not the compliance line.
  • Publish the definition as its own page, with exclusions listed and your name in the metric.
  • Deposit the extract somewhere with a persistent identifier, then stop describing that as proof of anything.

The uncomfortable summary is that provenance is a solved problem attached to the wrong end of the chain, and the field has adopted it because it is purchasable. Exposure to contradiction is not purchasable in the same tidy way; it has to be designed in, and it occasionally makes you publish a worse number than the one you started with. That is the cost of a figure that holds. Everything else is a well-signed assertion, and the statistics that get cited for years are never the well-signed ones — they are the ones somebody else has already tried and failed to knock down.

Leave a Reply

Your email address will not be published. Required fields are marked *

Verifiable Author Identity Previous post Verifiable Author Identity: sameAs, Credentials and Trust Chains
Trust Graph Next post The Trust Graph: How Engines Will Weight Verified Sources in 2027