Machine-Readable Brand

The Machine-Readable Brand: Building Your Owned Agent-Facing Layer

TL;DR

Everything in the owned agent-facing layer — summary files, agent manifests, structured data, feeds, endpoints — is a write. Every unit of value in it is created at a read you do not perform and cannot schedule.

Carriage is a chain, not a checklist: fetched × parsed × used. A zero at any link zeroes the product, and money buys the second and third links only.

A machine reads your file for exactly four reasons — permission, verification, capability, self-description. The first three are demanded by the reader. The fourth is offered by you, which is why it sits unread.

Demand arrives from the read side or not at all. Ads.txt went from roughly 7% of the top 500 publishers to more than 44% in three months, and nothing about the file changed — a buyer stopped paying for inventory that was not listed in one.

Your own server log settles the question in an afternoon, and it is the only scoreboard in AI visibility that nobody can sell you.

1. Publishing is a write; value is set at the read

The agent-readiness checklist circulating in 2026 is remarkably consistent from vendor to vendor. Ship an llms.txt, a plain-text file listing your best pages for language models. Publish an agent card, a JSON manifest saying what your site can do for an automated caller. Mark your pages up in JSON-LD, a structured-data syntax that states facts about a page in a form software can parse. Expose a feed or an endpoint. The pitch underneath the list is the part that actually sells it: unlike rankings, unlike citations, unlike whatever an engine decided about you last Tuesday, this layer is yours. You own the files. You control the facts.

The first half of that is true and the second half does not follow from it. You own the file, and ownership of a file is ownership of a write. Every unit of value in the owned layer is created at a read — performed by a crawler, an assistant or an integration, on a schedule set by someone whose costs you never see. The read is scarce, expensive to the reader, and not yours at all. Most of the money spent on machine-readable infrastructure in 2026 went to the half of the transaction that was never the constraint, and the wider AI-search visibility playbook has the same shape: the artefact is easy to produce and the reading of it is not.

The evidence is unusually clean because these files sit at fixed paths, so anyone with server logs can count who asked for them. Originality.ai tracked llms.txt across three million sites and found 4,088 files in June 2025 growing to 36,120 by May 2026. SE Ranking put adoption at 10.13% across 300,000 domains — and noted that among the fifty most AI-cited domains, exactly one had the file. Then Ahrefs looked at the read side across 137,210 domains with measurable traffic and found that in May 2026, 97% of llms.txt files received zero requests. Of the 3% that were fetched, SEO audit tools accounted for 21.7% of requests, GPTBot for 4.51%, ClaudeBot for 0.80%, and AI retrieval bots in aggregate for 1.1%.

Two independent datasets agree. Limy monitored more than 500 million AI bot events over 90 days and counted 408 requests targeting the file; MaxAEO examined nineteen customer log sets from February to April 2026, all from sites publishing it, and found 41 requests against roughly 1.1 million AI-crawler page fetches. The single most informative line in any of this comes from the Ahrefs study, and it is easy to skim past: AI bots never requested llms.txt on domains where the file did not exist. Cheap probing is exactly what a system that wanted the file would do.

The pattern is not confined to one format. In July 2026 API Evangelist fetched the agent-card paths of every host in the APIs.io catalogue — 20,185 that answered — and found 65 serving a card, of which ten passed every structural check in the A2A 1.0.0 specification. That is 0.29% of the most agent-literate population on the web.

What is the machine-readable brand?

It is the set of facts about your organisation that software can retrieve and act on without a human interpreting a web page — prices, identifiers, specifications, availability, permissions and capabilities, published in formats a machine can parse. It is not a description of your brand; it is the subset of your claims that a stranger’s system can use. That distinction does most of the work in this article, and it is what separates AI-readable feeds and endpoints from a well-formatted summary of yourself.

None of this makes the owned layer worthless. It makes it a layer where the thing you build is not the thing that pays: what pays is a fetch, and a fetch is a decision taken by a counterparty with its own budget. Designing for a decision you do not make is, as with every other form of earned visibility, the whole discipline.

2. The read chain: fetched, parsed, used

The industry argues about these artefacts as though the question were binary — does schema work, does llms.txt work — when each fails at a different link, for a different reason, and the fixes do not transfer. Treat every artefact as a chain of three.

THE READ CHAIN

C = f × p × u  — the carriage of a fact, per engine, per window.

f — the probability the artefact is fetched by an engine you care about. Observable in your own logs.

p — the probability its content survives the parse once fetched. Pipeline-specific, and not the same for a training crawl as for a live retrieval.

u — the probability the parsed content is used in the answer rather than discarded as unverifiable.

The expected value of publishing a fact is C × V. Because the terms multiply, a zero anywhere zeroes the product, and a strong artefact with f = 0 is worth precisely as much as no artefact at all.

Refusal one: nobody fetches it

This is the llms.txt result, and it needs no further arithmetic. When 97% of files receive no requests in a month and the retrieval bots do not probe for the path on sites that lack it, f is not small — it is zero, and no amount of curation inside the file moves a term that is set outside it.

Refusal two: fetched, then dropped

Structured data fails at the second link, and only in some pipelines. A 2026 searchVIU experiment across ChatGPT, Claude, Perplexity, Gemini and Google AI Mode found that during direct retrieval — the live fetch an assistant performs while answering — every system extracted only visible HTML, and JSON-LD, hidden Microdata and hidden RDFa were all ignored. Yet Volpini and colleagues, in an arXiv paper published in March 2026, measured a 29.6% retrieval-accuracy improvement for JSON-LD-marked content over plain HTML in standard retrieval-augmented pipelines, the architecture in which a system searches a corpus and feeds what it finds to a model.

Both results are correct, and the apparent contradiction is the point: p is a property of the pipeline, not of the markup. Ask which pipeline is reading you before deciding what a format is worth. The same page can be excellent input to an indexing crawl, valuable to a corpus builder, and invisible to the assistant that fetched it live thirty seconds ago — which is why the behaviour of deep-research modes and of ordinary AI Overviews diverge so often on the same URL.

Refusal three: read, and not used

The third link is the one nobody wants to price. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 and matched them against 4,000 control pages. AI Overviews citations moved by −4.6%, Google AI Mode by +2.4%, ChatGPT by +2.2% — all statistically indistinguishable from nothing. The study carried a real constraint worth stating: every page in it already had a hundred or more AI Overview citations before the experiment began, so this is a test on pages already inside the citation pool, not on invisible ones. Read carefully, it says the markup was fetched and parsed and then changed no decision. That is a u failure, and it is a different problem from an f failure, with a different remedy.

Why the f term is not for sale

Fetches cost the reader money, and the reader is already spending badly. Rebuilding Cloudflare Radar’s rolling 28-day window ending 21 July 2026 against its own 500-site panel, SEOmator put Mistral’s crawler at 3,389 pages fetched per referral sent back, ClaudeBot at 2,237:1, GPTBot at 217:1 and DuckDuckGo at 2.5:1. Those ratios are tied to their window and should never be averaged across windows, but the shape behind them is the shape behind the zero-click traffic model. Cloudflare’s May 2026 breakdown adds the reason: search-purpose crawling, the kind that can produce a citation with a link, made up under 10% of AI crawler requests. On 3 June 2026 Matthew Prince published Radar data showing automated requests had passed humans for the first time, at 57.5% of HTML traffic.

A reader with that cost structure spends fetches where they pay. It will not add a speculative path to its crawl plan because a publisher would like it to, any more than it will widen the source set behind an AI Mode answer out of courtesy. You can raise p by changing where a fact sits on the page. You can raise u by publishing facts that are worth relying on. You cannot buy f, and any strategy whose first move requires f to rise is a request, not a plan. The aggregate visibility figures hide this because they average over artefacts whose failure modes are unrelated.

Key takeaway

Diagnose before you spend. Zero fetches is an f problem and curation will not touch it. Fetches without citations is a p or u problem, and both of those are inside your control. Treating all three as “the site is not agent-ready” guarantees the budget lands in the wrong place.

3. Why a machine reads a file at all

There are exactly four reasons a machine fetches a file you published, and sorting your artefacts into them predicts the fetch data better than any argument about formats.

ReasonTypical artefactWho initiates the fetchEvidence it happens
Permissionrobots.txtThe crawler, before it actsUniversally honoured by major crawlers; Amazon documents that Amazonbot will fall back to a cached copy up to 30 days old if the live fetch fails
Verificationads.txt, app-ads.txtThe buyer, to release moneyAdoption among the top 5,000 programmatic sellers went from 8.5% in September 2017 to 51% by the end of February 2018 (Pixalate)
CapabilityMCP endpoint, feed, APIA user’s own tool, on demandDocumentation sites serving coding assistants are the one population where these files are demonstrably read
Self-descriptionllms.txt, curated summary, most agent cardsNobody97% of llms.txt files got zero requests in May 2026 (Ahrefs); 0.29% of reachable API hosts served an agent card (API Evangelist, July 2026)

The first three rows describe a reader who needs something: a permission it must have before acting, a fact it must check before paying, a capability it wants to invoke. The fourth describes a publisher who has something to say. That asymmetry is the entire result. The fetch belongs to the reader, and readers fetch for their own reasons.

There is a deeper reason self-description fails, and it is structural rather than a matter of adoption timing. A retrieval system exists precisely so that it does not have to take your word for what matters about you. Self-declared metadata that no user sees is trivially manipulable, so a system that ranked sources on it would be inviting exactly the spam it was built to filter — the same reasoning that killed the keywords meta tag two decades ago. Google’s own staff made the comparison: in July 2025 Gary Illyes confirmed Google does not support llms.txt, and John Mueller likened it to the keywords tag. This is not a gap waiting to be closed. It is a system doing its job, and it is why the factors behind AI product recommendations lean so heavily on things stated somewhere other than your own marketing.

So does llms.txt do anything?

For discovery-driven reads by answer engines, the 2026 log evidence says no — and the spec’s own August 2026 revision remains a navigation aid that grants and revokes nothing. Where it demonstrably earns its keep is the capability row: a developer points a coding assistant at your documentation and the file is fetched because a user configured it, not because a crawler discovered it.

The proof runs both ways

Ads.txt is the demand-side case. In August 2017 roughly 7% of the top 500 publishers had a file. From the end of October 2017, Google’s Display & Video 360 stopped buying inventory from sellers not listed in a publisher’s ads.txt where a file existed, and major buyers — Digitas, Omnicom’s Resolution Media — announced they would buy only from publishers who had one. By November 2017 more than 44% of publishers had shipped it; among the top 5,000 programmatic sellers, adoption ran from 8.5% in September 2017 to 51% by the end of February 2018. Nothing about the format changed in those five months. A reader’s budget started depending on it.

Security.txt is the control. RFC 9116 was published in April 2022, making it an actual IETF standard, endorsed by CISA and the UK’s NCSC — every institutional advantage llms.txt lacks. A 2026 scan of 240 million domains found growth from 82,000 domains in 2021 to 573,000 in 2026: a sevenfold rise, and still under 0.25% of all domains, with most of that increase attributable to hosting platforms auto-provisioning the file rather than organisations deciding to. Four years, an RFC number and two national cyber agencies produce a quarter of one percent when no reader’s budget depends on the outcome. Standardisation does not create demand; it only makes demand cheaper to satisfy once it exists — which is also the honest reading of what crawler permissions and content licensing have achieved so far.

4. The trust term is checkability

If f is set by the reader and p by the pipeline, u is the term you actually own — and it is governed by one property. A machine uses a statement when the statement can be checked against something you do not control. Everything else is decoration that costs parse budget.

A price in Product markup that matches the price at checkout is checkable, and cheaply: the system can fetch the checkout. A company registration number is checkable against Companies House. A model number with a stated capacity is checkable against a distributor listing and a spec sheet. In each case the cost of verification to the reader is one more request, and the payoff is a fact it can repeat without exposure — which is exactly the calculation behind entity authority as engines currently measure it.

Now the other column. “Market-leading.” “Trusted by thousands of professionals.” “Our most comprehensive guide.” There is nothing to check these against, so there is nothing to gain from believing them and a reputational cost to repeating them. They are not weak signals; they are unusable inputs, and a retrieval system’s handling of them is not a bug to be worked around. For example: replacing “industry-leading warranty” with “18-month warranty on parts and labour” changes nothing about your prose and everything about the statement’s status, because the second version can be contradicted — by your own warranty page, by a returns policy, by a distributor. A claim that can be contradicted is a claim that can be relied on.

This is also where the provenance work of the last two years earns its place. Content credentials and C2PA signing matter here not because engines currently rank on them but because they belong to the checkable column by construction: a signature is a statement about origin that a reader can test without trusting the signer. The same logic applies in reverse to AI content labelling — a label is worth something only where the labelling can be verified against something else.

Key takeaway

Audit your markup for unfalsifiable claims and delete them. Every property in a structured-data block should be a statement a stranger’s system could disprove in one request. If it could not, it is occupying budget in a pipeline that will discard it anyway.

5. The fetch proof: measure f yourself

Every number above came from someone else’s logs. Yours will differ, and the difference is the only figure that should drive your budget. This is the rare question in AI visibility you can settle without a tool, a vendor’s sampling frame or anyone’s permission: the raw access log records every request anyone made for anything you published.

THE FETCH PROOF

1. Window. Ninety days of raw access logs. Strip static assets; keep HTML and the specific paths of your machine-readable artefacts.

2. Verify before you count. A user agent is self-declared text. Do reverse DNS on the client IP, then forward DNS on the result, and confirm it resolves back. Cross-check against published ranges: OpenAI, Anthropic, Perplexity and Google publish machine-readable IP lists; Meta, Apple and Amazon publish documentation pages only.

3. Group by path × verified agent. Compute f for each artefact, for each engine that matters to you, and write down the window dates beside it.

4. Split the classes. Training crawlers and live retrieval fetchers behave differently and mean different things. The user-initiated agents — ChatGPT-User, Claude-User, Perplexity-User — are the ones whose fetch can end in a citation with a link.

Step two is not pedantry. LumenGEO’s 2026 crawler audit found that 2.0% of requests claiming a checkable identity came from outside that company’s published ranges, and the highest failure rate among checkable bots was GPTBot at 10.8% — seventeen of 158. The most famous identities are the most impersonated, which is the expected result. Report that number as a floor rather than a total: unverifiable is not a synonym for fake, and the honest audit has three buckets, not two. Web Bot Auth adds cryptographic proof per request, but the signed share is still small enough that reverse DNS remains the workhorse. Any technical SEO workflow that already parses logs can absorb this in an afternoon.

Reading the result

Zero verified retrieval fetches in ninety days means f = 0, and no editing inside the file changes a term set outside it. If most requests come from SEO crawlers, you are watching the industry study itself — in the Ahrefs data, audit tools made 21.7% of all llms.txt requests. And if pages are fetched heavily but not cited, you have ruled out the f problem, which is good news: p and u are the terms you can work on. Pair the log with a citation check so you are reading both ends; a SERP-less audit and a crawl log answer different halves of the same question.

Two limits worth stating plainly. A log proves a fetch, never a use: it cannot tell you whether what came back influenced an answer, only that something asked for it. And rendering matters before any of this: if your facts are assembled in the browser, a JavaScript-dependent crawl may record a successful fetch of a page that contained almost nothing, and your f term will look healthy while your p term is quietly zero. The same trap catches teams evaluating AI browsers, where the fetch is performed by something between a crawler and a user.

6. A worked example: Sedgemoor Instruments

Sedgemoor Instruments is a Somerset manufacturer of laboratory balances and calibration equipment: 41 staff, about £9.6m turnover, fourteen distributors across the UK and EU. The facts that matter commercially are unglamorous and specific — model numbers, capacities, readability in milligrams, calibration accreditation, lead times, territory coverage. Exactly what a buyer now asks an assistant rather than a sales desk.

In September the marketing lead ran ninety days of logs covering June to August. Five artefacts were in scope: an llms.txt published in February, hand-written over four hours and maintained for about an hour a month; JSON-LD on 340 product pages; an RSS feed; a public endpoint returning prices and lead times, built for two distributors; and robots.txt.

The counts, after verification: llms.txt received 61 requests in ninety days — 44 from three SEO suites, nine from a commercial data aggregator, eight unverifiable, and zero from any verified retrieval agent. Product pages received about 214,000 verified AI-crawler fetches. The distributor endpoint received 1,180 fetches, every one from two verified partner ranges. Robots.txt was fetched 9,400 times. That is the four reasons repeating themselves in one company’s logs: the permission file and the capability endpoint read constantly, the pages read at scale, and the self-description file read by tools that sell audits.

The October decision was not to delete the file. It was to stop curating it: llms.txt is now generated from the product database during the build, which takes its maintenance cost from thirteen hours a year to zero. The thirteen hours went into two changes with a defensible mechanism. First, the specification table — capacity, readability, calibration interval — was moved into the visible page body rather than living only in markup, on the searchVIU finding that live retrieval reads what is visible. Second, each product’s markup now carries its calibration certificate number, a fact a buyer’s agent can check against an accreditation register rather than take on trust.

The January re-audit showed spec answers appearing in two engines — suggestive rather than proof, since the sample is one company over one quarter and several things changed at once. The unambiguous number came from elsewhere. Generating both surfaces from one store exposed 41 product pages where the lead time in the markup disagreed with the lead time on the page. Those had been shipping contradictions to every reader for months, and nobody could have found them by reading either surface alone. It is worth being clear about which of the two outcomes is the reliable one: the citations are an anecdote, the 41 contradictions are a measurement.

7. One writer, many surfaces

Every additional machine-readable surface is another copy of the same facts, and copies drift. A price changes, a lead time slips, a certification lapses, an editor updates the page and not the feed. This is the failure mode that actually bites in the owned layer, and it is worse than the absence it replaced.

Worse, specifically, because checkability is the property that earned the trust term in the first place. An absent fact costs you a fact. A contradicted fact costs you the premise that your published facts can be relied on, and a reader that catches one disagreement has a cheap reason to discount the rest. You publish one estate, and its worst member sets the price.

THE SINGLE-WRITER RULE

Run only as many machine-readable surfaces as one build step can emit from one store.

Every surface generated, never authored. If a human has to remember to update it, it is already drifting.

Diff every shared fact across surfaces at build time, and fail the deploy on disagreement — the contradiction is the defect, not the deploy.

The seam: anything that changes without an editor touching it — prices, availability, identifiers, dates, lead times, certifications — belongs in the store. Prose belongs on the page.

The security.txt data shows what generation without a source of truth produces. Most of its adoption growth came from platforms auto-provisioning the file, and when a file exists there is roughly a 60% chance it points to a platform’s generic contact rather than anyone who can act on the report. Present and wrong is a real state, and it is reached by generating from the wrong store rather than by neglect.

This is also the test for how much to build. The right number of surfaces is not the number in the vendor checklist; it is the number your build can keep identical without human attention. For most mid-market operations that is two or three. A team that publishes a genuine interactive asset and keeps its figures in one place runs a smaller estate more honestly than a team with seven surfaces synchronised by memory, and the strategy layer should be sized the same way.

Key takeaway

Count your machine-readable surfaces, then count the ones a human updates by hand. The second number is your drift exposure, and the fastest safe reduction is usually generation rather than deletion.

8. The strongest objection: it is a cheap option

The best argument against everything above does not dispute the fetch data. It says the fetch data is irrelevant to the decision. Robots.txt had no readers before crawlers existed. Ads.txt had almost none in August 2017. A file costs twenty minutes to write and pennies a year to serve, and the payoff if a major engine ever conditions its behaviour on one is large and asymmetric. On that reading, publishing is buying a cheap option, and an option is not refuted by the observation that it is out of the money today. Nearly every practitioner who has actually read the log evidence still recommends shipping the file, for precisely this reason. It is the correct argument, and it deserves better than a repetition of the adoption numbers.

Here is the answer: they are right about the file and wrong about the cost. Ship it. Three costs sit behind it, and none of them is twenty minutes.

The measurement error is the expensive one. Publication gets booked as progress, and progress ends inquiry. Sedgemoor’s file cost four hours to write and thirteen hours a year to maintain, which is the real shape of the bill — the twenty minutes in the argument is the first draft, and curation is where the budget actually goes. A team that has shipped the file has, in its own reporting, done the agent-readiness work.

The second cost is divergence. A hand-maintained surface is a liability that grows with age, and an option that decays into a contradiction was never free. This is why the concession is specific: hold the option at genuine zero by generating the file, and spend nothing on curating it.

The third is watching the wrong thing. If demand arrives from the read side, the trigger to monitor is a read-side event, not a publishing milestone. In October 2017 the trigger was a buyer that stopped paying for unlisted inventory — not a working group, not a spec revision, not an adoption percentage. The equivalent triggers today are observable and none of them has fired: verified retrieval agents probing for a path on sites that do not publish it; an engine stating that a self-declared file conditions retrieval; a commercial counterparty refusing to transact without a machine-readable surface.

What would prove this wrong

If verified retrieval agents begin requesting these paths on domains that do not serve them, the non-probing finding inverts and self-description moves from the fourth row of the table into the second. If a major engine publishes that it fetches and conditions answers on a self-declared summary, the same thing happens faster. Either would make this analysis obsolete rather than merely incomplete, and both would appear in ordinary server logs before they appeared in any guidance. There is also a standing limit on the claim: where the reader is a user’s own tool pointed at your endpoint — a coding assistant loading documentation, a distributor’s integration, an assistant working through a multi-turn query chain with your endpoint configured — f is set by a configuration decision rather than by a crawler’s budget. That is a different and considerably better game, and it is the one worth building for.

9. What to do on Monday

None of this requires a rebuild, and most of it is measurement you have already paid for and never read.

  1. Pull ninety days of raw access logs. Strip static assets, keep HTML and the exact paths of every machine-readable artefact you run.
  2. Verify agents before counting anything: reverse DNS then forward DNS, cross-checked against published IP ranges. Record the unverifiable share as its own bucket.
  3. Compute f per artefact for the three or four engines you actually care about, and write the window dates next to every figure so the numbers stay comparable next quarter.
  4. Sort every artefact into permission, verification, capability or self-description. Anything landing in the fourth bucket gets generated and never curated.
  5. Move any fact you want quoted into the visible page body as well as the markup, on the evidence that live retrieval reads what is visible.
  6. Delete unfalsifiable claims from structured data — keep the properties a stranger’s system could disprove in one request.
  7. Add a build-time diff across every surface that repeats a fact, and fail the deploy on disagreement.
  8. Diary the trigger rather than the milestone: re-run the audit quarterly and watch specifically for probing on paths you do not serve.

The owned layer is worth building. It is simply not owned in the way the phrase suggests: you own the write, the reader owns the read, and everything useful in this discipline follows from designing for a decision taken on someone else’s budget. Publish the facts a stranger can check, generate every surface from one store, and let your own logs — not a vendor’s readiness score and not any tool’s dashboard — tell you which of them anybody actually reads.

Leave a Reply

Your email address will not be published. Required fields are marked *

Governance Checklist Previous post The 2027 AI-Era Link Building Governance Checklist