Original Interviews and Primary Sources

Original Interviews and Primary Sources: The 2027 Link-Earning Format

TL;DR

An interview is not a research method. It is a publication event: the moment a source with no URL becomes a document with one. Its value therefore tracks the retrieval gap it closes, not the standing of the person quoted.

Two filters sit between what a source knows and a sentence an engine can cite. The first is the container: the most primary records in Britain are the least reachable, because a crawler cannot submit a search form or read a photograph of a signed page. The second is the approval round, which removes precisely the properties that make a sentence citable.

Use the Retrieval Barrier to sort where the gap is, and the Prior-Clearance Brief to run interviews that survive a press office. The first four barriers pay well and expire; the last two pay slowly and do not.

1. The advice everyone is running, and the paper underneath it

The 2026 consensus is unusually tidy. Backlinko’s digital PR guide, updated in July 2026, states that expert commentary acts as a trust signal that makes content more cite-worthy. Demand Local’s agency playbook from April 2026 says a named expert quote inside a trusted publication is exactly what large language models extract. Involve Digital’s guide, from the same month, is the most explicit: the key asset for journalist outreach in 2026 is a spokesperson profile with verifiable credentials and pre-written quotes ready to deploy in tight turnaround windows.

Hold that last phrase. We will come back to it, because it quietly describes the failure mode of the entire genre.

The advice is not invented. It traces to one real, peer-reviewed result: GEO: Generative Engine Optimization, by Aggarwal and colleagues at Princeton, IIT Delhi, Georgia Tech and the Allen Institute, presented at KDD 2024. The paper tested nine content modifications across roughly ten thousand queries and found that three of them moved visibility materially: citing sources, adding statistics, and adding quotations. Reported lifts run between twenty-two and forty-one per cent depending on which metric you read, and keyword stuffing did nothing at all. It remains the only large-scale academic study anyone in this field can point to, which is why every playbook cites it and almost none of them read past the abstract.

What the GEO paper actually tested

Three details change the meaning of the result. The experiments ran on GEO-BENCH, a synthetic benchmark, against a generative engine the authors built to mimic Bing Chat rather than against live products. The manipulation was adding quotation marks and references to a page the researchers controlled. And the outcome variable was the visibility of that page — not whether the quoted person, or their employer, was named anywhere in the answer.

So the finding is about page formatting. Quotation addition is a typographic and structural change you can make to a draft in an afternoon, for free, without speaking to anybody. Conducting interviews is a supply-chain operation with scheduling, legal review and a per-unit cost measured in weeks. The profession read a result about the first and wrote a budget line for the second. That is not a small inferential leap. It is the difference between adding punctuation and hiring a newsroom.

The 2026 replication nobody quotes

There is now a follow-up, and it is unflattering. A 2026 paper introducing FeatGEO (arXiv 2604.19113) re-ran token-level GEO methods across three engines and found they failed to consistently improve citation visibility over an unmodified baseline. On one engine, visibility fell from a 13.34 per cent baseline to between 10.92 and 12.21 per cent under the standard heuristics. The authors’ conclusion is that isolated text-level edits are insufficient at scale and may disrupt the natural writing patterns models prefer to cite.

Two years of playbooks rest on a formatting effect that has not replicated. Meanwhile the sourcing programme built on top of it has real costs and, in most audits, thin results. Something in the chain from interview to citation is broken, and it is not the interview.

2. What an interview is, in retrieval terms

What is an original interview? In retrieval terms, an interview is the act of converting a source that has no address into a document that has one. Everything else about it — the craft, the rapport, the transcription — is production detail.

A person is the least retrievable primary source in existence. They hold an enormous amount of unpublished, first-order knowledge and they have no URL, no index membership and no fetchable representation of any kind. No amount of crawling reaches them. The only mechanism that moves what they know into the training corpus is somebody asking and then publishing the answer. That is the interview, and stated plainly it is a transcription event rather than a content format.

Once you see it that way, the valuation changes. The worth of an interview is not the standing of the person you spoke to. It is the size of the gap between what they know and what was already reachable. Interview a commentator whose views appear in forty published places and you have closed a gap of zero, at the cost of a fortnight. Interview the person who holds an unpublished operational record and you have moved something into the index that was not there before.

This is why the well-known-name heuristic misfires. Prominence is, definitionally, a measure of how much of somebody is already published. The more famous your source, the smaller the gap you are closing, and the more of the semantic credit their name absorbs when the answer is written. Anyone tracking entity-level authority signals has watched this happen: the quote gets carried, the professor gets named, the publisher gets a link if they are lucky.

What is a primary source to a retrieval system? Not the top of an evidential hierarchy. A retrieval system does not rank documents by their distance from the event. It fetches URLs that resolve, parses what comes back, and scores what it can read. Primacy is an epistemic property. Retrievability is an engineering one. They are different orderings, and in Britain they are close to inverted.

3. The legibility gradient: the best records are the worst documents

Run down the evidential ladder in any regulated British sector and watch the containers get worse as the evidence gets better. A trade magazine summary is clean HTML with a headline and a date. The industry report beneath it is a PDF. The statutory investigation beneath that is a PDF of a scanned signature. The register beneath that sits behind a search form. The standard beneath that is behind a paywall. The committee session that produced the standard was spoken aloud and never written down at all.

The field already knows the mechanic. Every AI-visibility guide published this year explains that crawlers operate as logged-out visitors sending plain GET requests, and cannot fill in forms, authenticate or read text rendered as an image. That fact has been applied almost exclusively to one thing: the reader’s own gated whitepaper. Ungate your PDF, the guides say, and be cited.

The same fact says something much larger, and nobody has drawn it. If a crawler cannot submit a form, then the entire public record that lives behind a search form is invisible too. That record is not yours to ungate. It is yours to restate.

THE FORM TEST

If you cannot send a colleague a single URL that opens the record directly — no search box, no session, no result set — then the record is not on the web. Neither is the fact inside it. A page that describes a database is not the database.

What this looks like in the British record

Coroners’ Prevention of Future Deaths reports are a clean illustration. They are public by design and central to institutional learning. A 2021 study in Journal of Patient Safety and Risk Management by Leary and colleagues examined 710 of them and reported the state of the archive without much diplomacy: the reports are published as PDFs, rarely as original digital documents, in many cases scans of signed pieces of paper and occasionally poor-quality photographs; they are not published in chronological order; they appear in batches; there is no centralised database, so it is not possible to know whether all reports are published or whether there are gaps. A separate scraping analysis found more than a thousand reports listing no recipients on the webpage at all, and six hundred recipient names entered with inconsistent punctuation.

Section 19 flood investigations are the same shape in a different sector. Under section 19 of the Flood and Water Management Act 2010, a lead local flood authority that investigates a significant flood must publish the results. That duty is discharged by roughly a hundred and fifty separate councils, each posting PDFs to its own site under its own naming convention, with no national register. Leicestershire’s page even lists investigations still being written, with estimated delivery dates. The statute says publish. Publishing is not the same as being on the web, and being on the web is not the same as being reachable.

4. The Retrieval Barrier

Six barriers stand between a public record and a citable sentence. They are not degrees of the same problem; each fails for a different reason and each is cleared by different work. Sort your sector’s records into these rows before you commission a single interview.

BarrierWhere it shows up in the British recordWhy retrieval failsWhat clearing it takes
1. Form-gatedPlanning portals, professional registers, tribunal decision searches, enforcement databasesThe crawler sends a GET request. The record only exists after a POST, so no URL was ever created for itOne stable, dated URL per record. The cheapest work in the set and the highest immediate yield
2. Container-gatedCoroners’ reports, Section 19 flood investigations, FOI responses, inquiry annexesThe text is a scan, a photograph or an unstructured blob. There is nothing sentence-shaped to extractTranscription into headed, dated HTML with the original document linked beside it
3. Scatter-gatedOne statutory report published separately by around 150 authorities, each to its own siteEvery fragment is individually fetchable and individually beneath notice. The pattern exists nowhereCompilation against a single schema. Slower, and the most defensible of the first four
4. Licence-gatedBritish and international standards, subscription market data, proprietary trade databasesReproduction is not yours to authorise, whatever the crawler can reachPublish facts about the document and never its text. Yield is capped by law, not by effort
5. EphemeralCommittee sessions, conference floors, site visits, technical panels, inspection walk-roundsThere is no record to fetch, because none was ever madeBe the party that makes the first record. Nothing can close this barrier except somebody writing
6. CustodialThe practitioner’s own head: what was decided, what failed, what the number excludesNo document exists and none is owed to anybodyAn interview — and then the approval round, which is section 5

Read the colours as this year’s yield per pound. Then read them backwards for durability, because the two rankings run in opposite directions, and that inversion is the whole of the planning problem. The barriers that pay best are the ones a publishing programme can close without asking you. The barriers that pay worst cannot be closed by anybody.

Key takeaway

Your sector’s most valuable evidence is usually public, free and unreachable. Sort it by barrier type, not by topic. Barriers 1 to 3 are legibility work — transcription, schemas and stable URLs — and are the cheapest link programme most firms have never costed. Barriers 5 and 6 need a person and a question.

5. The second filter: why approved quotes are unusable

Barrier six is the one the whole genre is aimed at, and it has a property no other citable asset has. An interview is the only research output with a living counterparty — somebody who retains rights over the sentence after you publish it, can withdraw it, can give the same material to your competitor next week, and is named alongside you in every downstream restatement.

Everything strange about the format follows from that. The most consequential is the approval round. Quote approval has been documented in journalism for well over a decade: Jeremy Peters described the practice in the New York Times in 2012, reporting that quotations came back redacted and stripped of anything mildly provocative. The paper subsequently forbade after-the-fact approval; the Washington Post issued ethics guidance after a reporter shared a full draft with officials. Journalism treats approval as a threat to independence. Marketing treats it as a scheduling step.

Both are missing the mechanical point, which is this. Quote approval is a filter that runs in the opposite direction to citation value. A communications review is not lazy or malicious. It is doing its job, and its job is to remove downside. So it removes unhedged numbers. It removes denominators, because a denominator invites a rate. It removes admissions. It removes anything naming a competitor. It removes any figure the legal team cannot defend if a customer quotes it back.

Now list the properties that make a sentence retrievable and worth citing: specific, dated, scoped, checkable, falsifiable, carrying a number with a denominator attached. It is the same list. The approval round is not a diluted version of the interview — it is the interview run backwards. You are paying to create a quotable sentence while a second party works, entirely in good faith, to make it unquotable.

What survives is directional generality: pressure will increase, expectations are rising, the market is maturing. That material is not merely weak. It is the exact register a language model produces for nothing, on demand, at any length. An expert quote that has survived a comms review is a press release with a person’s name on it.

Which returns us to that phrase from the April 2026 playbook: a spokesperson profile with pre-written quotes ready to deploy. Read literally, that is a recommendation to manufacture the quote before the question exists — approval-cleared, context-free, and by construction free of any particular. The industry has industrialised the filter and filed it under preparedness. BuzzStream’s State of Digital PR 2026 reports the predictable consequence from the other side: falling trust in expert commentary, driven by fake experts and by real experts’ names being attached to machine-assisted copy.

The filter’s strength tracks employer size

This gives an unglamorous but reliable sorting rule. The approval filter is a function of how many people are paid to manage reputational risk. A FTSE-listed operator has a communications function, a legal function and a policy function, and a quote passes all three. A sole practitioner has none. So do academics, retired officeholders, trade-body volunteers, standards panel members, and the technical staff of small regulators and inspectorates.

The citable expert is usually the one with the least impressive employer. That runs against every instinct in media relations, where the logo is the point, and it is the single highest-leverage change most interview programmes can make. It is also the honest fallback for sectors with no statutory record at all, which we come to in section 8.

6. The Prior-Clearance Brief

You cannot abolish the approval round. You can route around it, by never asking for a new sentence in the first place. The insight is that most sources have already said something under a regime far more consequential than your article, and that material has already been cleared by people more cautious than any press officer.

THE PRIOR-CLEARANCE BRIEF

Step 1. Before making contact, list every document the source has already put their name to under a regime of consequence: a regulatory filing, a submission to a consultation, evidence to a committee, a standards panel note, a peer-reviewed paper, a professional body response, a signed statutory report. Their name is on it. The clearance already happened, at a higher bar than yours.

Step 2. Write every question as a request to explain a sentence they have already signed, not to produce a new one. Not what is your view on surface water pressure, but your report attributes this event to highway drainage capacity — what does that attribution exclude?

Step 3. Ask for the decomposition, never the headline. The denominator, the period, the exclusion, the cases that did not qualify. Arithmetic performed on already-cleared material is usually still cleared, and it is where every citable particular lives.

Step 4. Never ask for a forecast, and never ask for a comparison with a named competitor. These are the two categories a review removes without discussion, and including them poisons the whole round: one flagged question routes the entire draft to legal.

Step 5. Send it back as a fact check, not an approval. Please confirm these figures are accurate is a question the source can answer alone. Please approve this quote is a question that routes to communications. Same document, different department, different outcome.

The test on the finished piece: if it would still make sense with the source’s name removed, you did not run an interview. You commissioned a paragraph.

This is a narrower instrument than it looks. It does not produce colour, personality or narrative, and it is useless for the sort of profile piece that earns coverage on charm. What it produces is particulars attached to a name that a third party has already accepted responsibility for — which is what survives paraphrase into an answer, and what a citation-recovery audit is looking for when it asks why a page was fetched and then dropped.

7. Worked example: Penrose Water Engineering

Penrose Water Engineering is a flood-risk and drainage consultancy in Exeter with 38 staff and £6.4M in fee income. It sells design, modelling and expert-witness work, and it does not sell remediation, so it has no product to defend. Its buyers are local authorities, housebuilders, insurers and litigation solicitors.

Round one: the expert interview programme

Between March and August 2025 Penrose spent £24,000 on eighteen interviews, published as a Q&A series. The list was strong on paper: a water company’s head of wastewater networks, two university hydrologists, a trade body director, four lead local flood authority officers, ten consultants. It produced 34 published quotes.

Twenty-one of the 34 were direction-of-travel statements. Eleven of the eighteen interviews went through a press office, with a median of nineteen days from draft to sign-off, and the approved versions had removed every figure that carried a denominator. A September 2025 probe across forty buyer-intent prompts found Penrose named in two — and both traced to the single interviewee with no employer communications function, a retired drainage engineer who had chaired a standards panel.

Round two: going at the record instead of the people

From September 2025 Penrose spent £58,000 over nine months on a different operation. It compiled 411 Section 19 flood investigation reports published between 2019 and 2025 by 96 lead local flood authorities into one dated HTML record per report, each with a stable URL, a fixed schema — date, authority, watercourse, properties affected internally, mechanism, risk management authorities named, recommendations — and the original council PDF linked beside it. Seven months of that was document handling. Nothing about it was creative.

Only then did it run nine interviews, with LLFA — lead local flood authority — officers, entirely on the Prior-Clearance Brief. Every question asked what a sentence in that officer’s own published report meant, excluded or assumed.

By June 2026 Penrose appeared in 23 of the same forty prompts and held 31 referring domains: two insurers, a university engineering department, a water company’s own consultation response, three planning consultancies, a parliamentary written submission — and, decisively, four council pages linking to Penrose’s version of their own report, because it was easier to link to than their PDF. The nine interviews yielded six usable quotes in total, but every one carried a number that appeared nowhere in the officer’s published report.

Four results that did not flatter the programme

First, there was no moat. A competitor scraped the compilation within eleven weeks and republished a thinner version. What Penrose held was priority and inbound links, not exclusivity. This is a rent on a temporary legibility gap, and budgeting it as an asset is a category error.

Second, they inherited a maintenance liability they had not created. Two authorities asked for changes. One report had been superseded by a later investigation Penrose missed. A third council withdrew its PDF while the transcription stayed up. They now carry a correction duty for records they do not own, and one wrong transcription is materially worse than no transcription at all — the failure mode that provenance and content credentials exist to make visible.

Third, the Prior-Clearance Brief reduced the filter without removing it. Five of the nine officer interviews still routed to a communications team, two took over thirty days, and one authority declined outright on the ground that commenting on its own published report might prejudice a live claim.

Fourth, the gains were lopsided by engine. Most of the movement landed in two of the five surfaces tracked; a third barely shifted, because the compilation sat outside the candidate pool it draws from. Anyone reading engine-level variance across a multi-turn research session will recognise the shape: presence is not portable, and a single blended visibility number hides which surface actually moved.

8. The strongest objection: you are doing the state’s product development

Here is the hardest version of the counter-argument, and it should be conceded flatly rather than managed. Every barrier in rows one to four is a defect the record’s owner has an incentive to fix, and the owners have started fixing them. The moment they ship, your position closes.

The evidence is unambiguous. From January 2023 the judiciary began publishing the full text of Prevention of Future Deaths reports directly on webpages rather than as attachments, explicitly to improve accessibility — retiring a large part of exactly this arbitrage in a single release. In 2026 MHCLG moved flagship datasets to CSV on the Web, pairing plain files with machine-readable metadata, and said openly that it was doing so because AI-driven reuse is coming. data.gov.uk is being rebuilt as the National Data Library. GOV.UK already runs a Content API that serves published content as structured JSON, and its publishing schemas enforce a consistent shape across every content type.

When a barrier closes, the compiler’s retrieval position closes with it. The referring domains earned during the window do not unwind — an earned link is banked — but citation share is re-competed on every query. Assume the position is temporary.

Four bounds on the objection

One. The counter-argument’s own documents undercut its timeline. The National Data Library roadmap, published June 2026, states that the original data.gov.uk approach failed, producing broken links and low usage — sixteen years after launch. Legibility programmes are slow because they are funding decisions, not technical ones.

Two. Central barriers close centrally. Barrier three does not, because closing it requires roughly a hundred and fifty separately funded authorities to agree a schema and then maintain it. That is a policy act, not a deployment, and it is why scatter is the most defensible of the first four rows.

Three. Barriers five and six cannot be closed by any publishing programme, because there is no record to publish. No API can serve what was said aloud in a technical panel and never minuted. This is the only permanently defensible position in the table, and it is the one the interview was always for.

Four, and honestly: everything above row five is a two-to-four-year rent. Priced as a rent it remains excellent value, because the returns arrive early and the inbound links persist. Priced as an asset it will surprise you twice — once when a rival copies it, and again when the issuing authority makes it redundant.

Key takeaway

Yield and durability run in opposite directions down the Retrieval Barrier. Run rows 1 to 3 for reach this year, on the explicit understanding that you are renting. Run rows 5 and 6 for tenure. A programme that only does one of the two is either short-lived or too slow to fund.

9. The objection with no clean answer

A large share of businesses have no statutory record at all. Consumer brands, most software, agencies, most professional services: no register, no inspectorate, no inquiry, nobody with a duty to write anything down about them. For those firms the table collapses to row six, the approval filter is unavoidable, and the honest nearest move is the employer-size rule from section 5 — interview practitioners whose employers have no communications function, accept a smaller and slower output, and stop pretending the difference is a craft problem.

There is a second boundary worth naming. Some organisations have a rich statutory record and it is about them — enforcement notices, upheld complaints, inspection scores. Making that legible is self-harm, and the correct move is to say so out loud in the planning meeting rather than discover it in week nine.

10. What this changes about link acquisition

Prospect for custodians, not commentators. The list you want is people obliged to produce or hold a record: scheme secretaries, registrars, inspectorate technical staff, standards panel chairs, LLFA officers, trade body committee members, ombudsman case officers. Sort by what they are required to produce, not by audience. No backlink tool generates this list, because custody is a legal property rather than a web one — the same reason it does not appear in any competitor-driven prospecting workflow built on link graphs.

Treat the issuing authority as a link prospect. When you make a body’s own record usable, that body acquires a live reason to point at you. The judiciary’s own guidance on coroners’ reports directs readers to a third-party Preventable Deaths Tracker built by researchers who did the compilation work. That is a public institution linking outward because an outsider’s version was more usable than its own. It is among the most defensible inbound links available in any regulated sector, and it is earned by document handling rather than by pitching.

Stop pitching the list everyone pitches. Muck Rack’s Generative Pulse analysis, reported in March 2026, found that the journalists most frequently pitched by PR professionals and those most frequently cited by AI engines overlap by roughly two per cent. A media list built on circulation is close to orthogonal to the citation set. The same research programme puts earned media at 84 per cent of AI citations across 25 million cited links, with journalism at 27 per cent — so the channel is right and the targeting is wrong.

Buy transcription, not commentary. The invoice for a legibility programme is document handling, schema design and hosting, plus a stable URL policy. It looks like technical publishing rather than media relations, and it is priced accordingly. It is also the only link programme where your counterparty had a statutory duty to produce the thing you are republishing, which removes the negotiation entirely.

One structural note for anyone weighing this against the usual portfolio of tactics. Legibility work produces pages other people have a reason to cite for years, which is a different asset class from placed guest content or inserted links in existing posts, and it behaves differently under scrutiny. There is nothing to disclose, because nothing was bought. That property is worth more each year, as the authenticity premium gets harder to fake and easier to check.

11. The Monday checklist

  • List every record class in your sector that somebody has a legal duty to produce. Statute, regulation, licence condition, professional code. Do not filter yet.
  • Apply the Form Test to each one. If you cannot send a colleague a URL that opens a single record, mark it barrier 1.
  • Sort the survivors into the six barrier rows. Count how many sit in rows 1 to 3. That number is your legibility budget for this year.
  • Pick one record class and transcribe twenty records. Fixed schema, one stable dated URL each, original document linked. Twenty is enough to see whether anyone fetches them.
  • Name the correction owner before you publish. Who fixes a wrong transcription, and inside what window. Do not skip this step; it is the one that ends badly.
  • Re-cut your interview list by employer size, smallest first. Move every source with a press office to the bottom.
  • Rewrite your next interview brief against the five Prior-Clearance steps. Delete every forecast question and every competitor comparison before you send it.
  • Send the draft back as a fact check, not an approval. Use the words please confirm these figures are accurate.
  • Set a forty-prompt probe now, before publishing, and re-run it monthly by engine rather than blended.
  • Diary a twelve-month review of every barrier 1 to 3 asset, and ask one question: has the issuing body shipped its own version yet?

The format the map calls the 2027 link-earning format is real, but it is not the Q&A page. It is the position you hold when you are the reachable version of something that mattered before you arrived. Interviews are how you get there when no record exists. Transcription is how you get there when one does, and nobody can read it.

Leave a Reply

Your email address will not be published. Required fields are marked *

Lived Proof Corroboration Previous post Lived Proof: Why Reddit-Style Corroboration Beats Polished Pages
Experience-Led Content Engine Next post Building an Experience-Led Content Engine That AI Rewards