TL;DR — Every AI content deal since 2023 has been described as a sale of content. None of them were. What a buyer pays for is a release: a covenant not to sue, a warranted supply, and occasionally exclusivity or a name worth borrowing. The court in Bartz v. Anthropic priced the two halves of the transaction separately — training on lawfully obtained books was fair use, obtaining them unlawfully was not — and the second half settled at roughly $3,000 a work. The reading was free. The taking cost $1.5 billion. An open web page is served, not taken, which is why most sites hold almost nothing releasable and should stop pricing their archive. What they can sell is supply. What they should take in payment is attribution, not only cash. Two instruments follow: the Four Saleable Goods, and the No-Deal Test.
The playbook the whole market is running
The advice is now standardised, and it is repeated with enough confidence that it has stopped being examined. Audit your archive. Quantify your crawl volume. Assemble a media kit showing depth, freshness and topical authority. Register with a licensing marketplace, publish machine-readable terms, set a per-crawl price, and wait for the money that the largest publishers have already been paid.
The evidence behind it is real and should be stated at full strength before it is taken apart. Tenpoint Labs counted 12 publicly announced AI content licensing deals in 2023 and 91 by the end of 2025, with a mid-2026 projection of 127. News Corp’s agreement with OpenAI is reported at $250 million over five years. The Wall Street Journal put Amazon’s deal with The New York Times at $20–$25 million a year — close to one per cent of the Times’s 2024 revenue — covering training plus summaries and excerpts in Alexa. Underneath the headline names, an infrastructure layer has assembled itself: TollBit reported more than 3,000 publisher clients, ProRata more than 500 signups by mid-2025, and Cloudflare — sitting in front of roughly a fifth of the web — resurrected HTTP 402 to let any domain owner charge per request.
So the market exists, the money is real, and the infrastructure is built. The playbook still contains one false sentence, and it is the first one: that what changes hands is content.
Nothing in a content licence is scarce. Text is non-rival, and the frontier labs are not short of prose. A model that has ingested a substantial fraction of the crawlable web does not become measurably better because one more competent 2,000-word explainer joins the pile. If you price your archive by its quality, its word count or its crawl volume, you are pricing an input the buyer already has in effectively unlimited supply. That is why so many publishers report the same experience: a media kit full of impressive numbers, met with either silence or an offer that reads as an insult.
The deals that clear at eight and nine figures are not paying for better text. They are paying for something adjacent to the text, and the adjacency determines whether you receive a cheque, pennies, or nothing at all.
What a court actually put a price on
The clearest pricing signal in this market came from a courtroom rather than a negotiation, and it has been consistently misread.
In Bartz v. Anthropic, Judge William Alsup split the question in two. Using lawfully acquired books to train a large language model was transformative fair use. Acquiring and retaining a library of more than seven million pirated books was not, and buying legitimate copies afterwards did not cure the earlier act. That exposure ran to statutory damages capped at $150,000 per work for wilful infringement. Rather than test it at trial, Anthropic settled. On 20 July 2026, Judge Araceli Martinez-Olguín entered final approval: $1.5 billion, covering a defined list of 482,460 works, at roughly $3,000 a work before fees, with destruction of the pirated dataset and no coverage of outputs.
Read the two halves against each other. The court valued the reading at zero and the taking at $1.5 billion. What had a price was never the use of the text. It was the provenance of the copy. Norton Rose Fulbright’s 2026 survey of the docket makes the same point in lawyerly terms: there was no exemption for an unlawful acquisition simply because the eventual use might be transformative.
That distinction is the hinge of every licensing conversation, because an ordinary web page is not taken. It is served — by your own infrastructure, on request, in response to a fetch you could have refused. Whatever else the fetch is, it is a poor foundation for the kind of claim that produced a $3,000-a-work settlement. The books case was expensive precisely because the copies came from shadow libraries. Your copies come from you.
Why “it is already in the weights” is a pricing fact, not a grievance
Publishers usually raise prior ingestion as an injustice. Treat it instead as an input to the price. A buyer weighing a licence is asking what changes if they sign, and for historical archive material on the open web the honest answer is often: their legal position improves slightly, and their product does not change at all. The English High Court sharpened this in Getty Images (US) Inc & Ors v Stability AI Ltd [2025] EWHC 2863 (Ch), handed down on 4 November 2025, where Joanna Smith J held that the model weights did not store reproductions of the training works — there were, in the court’s framing, no copies inside the model. An appeal was listed to be heard before December 2026, but the finding tells you what a licence over an already-ingested archive is worth to a buyer who has already shipped the model.
What is an AI content licence actually buying?
Not the content. A licence conveys permission, a warranty and an indemnity over material the buyer can usually obtain anyway, which is why price tracks the risk and operating cost you remove rather than the quality of what you wrote. The practical consequence is that two sites with identical editorial standards can be worth very different sums, and the difference has nothing to do with the writing.
The Four Saleable Goods
A content licence bundles up to four separable goods. They have different buyers, different prices and different owners, and most negotiations go badly because both sides argue about a single number covering all four. Separate them and the conversation becomes tractable.
| The good | What the buyer is actually purchasing | Who can credibly sell it | How it prices — and what it costs you |
| Release | A covenant not to sue, plus warranties and indemnity over past and future use — certainty on the balance sheet. | Anyone holding an enforceable, provable claim in a jurisdiction where the buyer has assets. | Priced off the discounted cost of the claim you give up. Costs you the claim itself, permanently. The only good with a court-tested benchmark. |
| Supply | A maintained, structured, warranted feed with defined freshness, schema stability and an availability commitment. | Anyone who can operate a feed to a service standard — including parties whose content is already free to all. | Priced as a subscription against the buyer’s cost of building and maintaining the same pipeline. Costs you engineering time, not rights. |
| Exclusivity | Denial of the same material to named rivals for a defined term and territory. | Anyone whose material has no close substitute in the buyer’s domain. | Priced as a premium on the other three. Costs you every other counterparty for the term — usually far more than it pays. |
| Endorsement | Your name attached to the buyer’s product surface, and a defensible provenance story for their users. | Anyone with a mark that reassures a sceptical audience. | Rarely priced in cash at all. Costs you reputational exposure to a product you do not control. |
Release: the only good with a court-tested price
Release is what News Corp sold. It is what the Times sold to Amazon while continuing to litigate against OpenAI and Microsoft, and it explains an arrangement that otherwise looks incoherent: license to one buyer, sue another. Litigation capacity is not a distraction from the licensing business. In 2026 it is the licensing business, because a party that can credibly and expensively sue is the only party whose signature removes something from a buyer’s risk register.
The market concentrates accordingly. The Media and the Machine deal database records news and journalism as the largest single block of disclosed agreements, and CASRAI’s 2026 review of scholarly publishing reaches the same conclusion from the other end of the market: value accrues to holders of large, well-cleared, long-running backlists with in-house rights teams. A learned society running two journals through a hosting partner is crawled just as thoroughly and extracts nothing, not because its research is weaker but because its release is worth less.
Supply: the good almost anyone can sell
The most instructive licensing business in the world belongs to an organisation with no release to sell at all. Wikimedia gives its content away under an open licence; nobody needs its permission for anything. In January 2026 it nonetheless formalised paid enterprise arrangements with Amazon, Meta, Microsoft, Perplexity and Mistral, converting disorderly crawler traffic into structured, paid, real-time access. Nothing about the rights changed. Everything about the delivery did. Buyers paid for schema stability, refresh guarantees, and someone accountable when the feed breaks at three in the morning.
That is the shape the market is moving into. Mark Riley’s tracking of 94 publicly announced deals found that only about four in ten now include training rights, down from near-universal inclusion in 2023 and the first half of 2024; the newer agreements — Getty with OpenAI, AP with OpenAI — were announced without them. Retrieval-augmented generation, where a model fetches live sources at answer time rather than relying on memorised text, changed what buyers need. They need today, delivered reliably, in a shape their pipeline can ingest. If you are deliberate about which AI training sources you feed, you are closer to a supply product than to a content archive, and should price accordingly.
Exclusivity and endorsement: the two you probably should not sell
Exclusivity looks like the premium term and usually functions as a trap. A licensing catalogue published by Presenc AI in April 2026 notes that some agreements carry partial exclusivity barring the publisher from licensing the same material to named competing labs. For a large rights-holder with one dominant buyer that can be rational. For everyone else it converts a non-rival asset into a rival one and hands the buyer a veto over your future revenue, in a market where the identity of the dominant buyer has changed roughly every eighteen months.
Endorsement is subtler. A buyer wants your mark on their surface because it makes their answers look trustworthy, and that is a real good with real value — but it is almost never paid for in cash, and it exposes you to a product whose failure modes you cannot see. The one thing Getty actually won on was narrow and historic trade mark infringement, where its watermark surfaced in generated images. The copyright claim collapsed; the mark travelled anyway. Marks move through these systems more readily than rights do, which is an argument for licensing your name deliberately rather than discovering it has been borrowed.
The No-Deal Test
Before you name a number, answer one question honestly. It sets your price band more accurately than any valuation model, and it is the question the buyer has already answered internally.
THE NO-DEAL TEST
What does the buyer do on Monday morning if you say no?
1. Nothing changes. They fetch what they need, face no meaningful claim, and their product is unaffected. Your price is zero — you are not negotiating a sale, you are negotiating over a gift. Ask for attribution instead.
2. They proceed wearing a risk. They can take it, but doing so leaves an exposure on the register. Your price is the discounted cost of that exposure: probability of a claim, times its size, times the buyer’s cost of capital.
3. They substitute, at a cost. A rival source exists but is worse, slower or messier. Your price is their switching cost — the quality gap plus the integration work — and nothing above it.
4. They cannot proceed. The product does not function without your material. Your price is a share of the product’s value, and you should be asking for a percentage, not a fee.
Most open-web publishers are in band one and negotiating as though they were in band four. Most operators of a genuinely proprietary dataset are in band three and never ask.
The test does something a valuation cannot: it tells you which currency to request. Bands three and four support cash. Band one supports credit — and credit, unlike cash, compounds. Bands two and three are where most real negotiations live, and both reward precision about what exactly you are conceding.
The United Kingdom discount, and why it is not permanent
Where you sit changes what band you are in, and British rights-holders are currently marked down. On 18 March 2026 the Government published its Report on Copyright and Artificial Intelligence with an accompanying impact assessment, discharging obligations under sections 135 and 136 of the Data (Use and Access) Act 2025. After more than 11,500 consultation responses it abandoned the previously preferred broad text and data mining exception — automated bulk analysis of copyright works — with an opt-out, and declined to endorse mandatory licensing either. The House of Lords Communications and Digital Committee had recommended a licensing-first approach on 6 March 2026 and urged ministers to rule the opt-out model out entirely. The result is the status quo, monitored: permission is generally still required, and the Creative Content Exchange pilot is the Government’s main practical intervention.
That sounds favourable to rights-holders until you set it beside Getty. The claim that mattered most there was abandoned mid-trial because Getty could not establish that training had occurred in the United Kingdom. A right you hold but cannot enforce against an offshore act of training is not a claim a buyer will pay much to extinguish. So a British publisher’s release good is discounted — not because the right is weak, but because the jurisdiction is. The discount is contingent: an appeal, a legislative move in 2027, or an enforcement mechanism attached to the Exchange would all reprice it, which is a strong reason to keep deal terms short. For anyone weighing cross-border link acquisition or already working across European markets, the same asset can sit in different bands in different jurisdictions at the same time.
How much is my content worth to an AI company?
It is worth the cost of the buyer’s next-best alternative, not a rate per word or per crawl. If they can obtain equivalent material freely and face no credible claim for doing so, the cash value is close to zero regardless of editorial quality — which is why the answer for most sites is to negotiate for attribution and supply terms rather than a fee.
Cash and attribution are different currencies
A licence can be settled in cash, in credit, or in both, and the field routinely assumes the first automatically produces the second. It does not. Digiday’s 2026 publisher reporting was explicit that partnership agreements do not guarantee citations: the commercial relationship and the product’s citation-selection behaviour are separate systems.
Goodie’s study of 31 million AI citations across eleven surfaces shows both edges of this. The Associated Press blocks every OpenAI crawler in its robots file and still draws roughly 80 per cent of its citations from ChatGPT, because licensed material reaches the model through the contract rather than the crawl. In the same sample, ChatGPT cited nytimes.com zero times and Gemini zero, while Grok produced 37,642 citations, Google’s AI Overviews 8,007, and Perplexity — a company the Times is suing — 4,303. A licence moves your visibility inside one buyer’s product and almost nowhere else. It is a distribution agreement with one distributor, and it should be valued like one.
This is why attribution deserves to be negotiated as consideration rather than accepted as courtesy. It costs the buyer almost nothing in cash and it is the only term that pays you in a currency with compounding returns. The Really Simple Licensing vocabulary, a machine-readable way to state licence terms, recognises this directly: alongside subscription, pay-per-crawl and pay-per-inference it defines a payment type called attribution, in which the credit is the price.
The displacement floor
Before accepting any figure, price what the licence costs you in visibility. The arithmetic is unglamorous and takes twenty minutes.
- Count the answers per month in which your material currently appears with your name attached, using whatever citation and entity-authority tracking you already run.
- Assign each a value — not a click value, a presence value: what you would pay a media buyer to be named in that answer for that query.
- Estimate what share survives the deal. An agreement conveying ingestion rights with no attribution obligation converts named appearances into anonymous ones.
- Multiply by the term, and treat the result as a floor beneath which no cash offer is acceptable.
A specialist publisher named in 900 answers a month, valuing presence at £2 each, is putting £21,600 a year of visibility on the table. An offer of £15,000 for unattributed ingestion is not a small deal. It is a loss with a cheque attached — and unlike a manual action you can recover from, a signed rights grant does not reverse when you change your mind.
Key takeaway
Cash is paid once and attribution accrues, so the term that looks like a courtesy is usually worth more over a three-year horizon than the term everybody argues about. Price the visibility you are giving up before you price the rights you are granting.
The clauses that decide which currency you get
Six terms determine whether an agreement complements the visibility you have earned or quietly replaces it, and buyers rarely resist them because none of the six costs much to concede.
Start with scope by purpose. Ingestion for training and retrieval for display are different grants with different consequences, and bundling them is the buyer’s convenience, not yours. Training rights are permanent in effect — a model cannot forget on request — while retrieval rights expire with the term. Grant retrieval readily; grant training slowly, separately, and for more.
Then term and survival. In a market where the dominant buyer changes every eighteen months and a British appeal judgment is pending, a five-year perpetual grant prices a legal environment that will not exist by year two. Twenty-four months with a renewal option is defensible; perpetuity almost never is. Insist that the grant terminates rather than survives, and that termination triggers deletion of the licensed corpus.
Third, the form of attribution, which is where most drafts are vague enough to be meaningless. Specify the surface, the wording, whether the credit is a live hyperlink or a plain-text name, and whether it appears in the answer body or a collapsed source panel. The difference between an inline named credit and a link buried in a source tray is the difference between a placement and a footnote — the same distinction that separates a citation from a link in ordinary search, and it deserves the same scrutiny here.
Fourth, reporting and audit. Ask for periodic usage data: how often the material was retrieved, and how often it was surfaced with credit. Buyers who intend to honour attribution rarely object; those who object have told you something. Fifth, a most-favoured-nation rate on comparable terms, which costs nothing today and protects you against the market repricing upward while you are locked in. Sixth, exclusivity carve-outs — if exclusivity is unavoidable, scope it narrowly by product surface and territory, never by content class.
Key takeaway
Split the grant by purpose, keep the term short, and write the attribution clause with the specificity you would apply to a payment schedule. Those three moves change the economics of a deal more than the headline number does.
Marketplaces, collectives, and the pennies problem
If a bilateral deal is out of reach — and for the overwhelming majority of sites it is — the marketplaces are the obvious fallback, and their arithmetic deserves stating plainly before anyone builds a forecast on it.
Cloudflare’s pay-per-crawl mechanism sets a flat per-request price and answers unpaid crawlers with a 402. Leaky Paywall ran the sums in 2026: on a site with a million monthly pageviews, AI crawlers might account for ten to twenty thousand hits, which at a tenth of a penny per page yields roughly twenty dollars a month and at a full penny roughly two hundred. That is a rounding error against payroll. A per-fetch toll charges for access, and access was never the scarce good.
The attribution-weighted models fare better for differentiated material. ProRata pays a 50/50 share weighted by how often a source contributes to an answer, which rewards distinctive material and penalises commodity explainers; TollBit prices per fetch and lets rights-holders keep the full amount, charging buyers a transaction fee. Microsoft’s Publisher Content Marketplace, announced in February 2026, follows a pay-per-use model with the take rate still undisclosed. The May 2026 Tow Center analysis reported by Nieman Lab warns about exactly that opacity, noting that publishers may find themselves negotiating with the same firms that operate the scaffolding, and using Spotify’s thirty per cent take as the benchmark against which these arrangements should be judged.
Presenc AI’s catalogue puts the certainty premium for bilateral deals over marketplace participation at roughly two to ten times at the per-citation level. That gap is not a quality premium but the price of certainty, which is the same good the release delivers — evidence for the thesis rather than against it. Collective licensing is the structural answer for the long tail, and the RSL Collective is the most credible attempt at it, but pay-per-inference requires buyers to log which sources shaped which output, and few have built that. Joining is free and non-exclusive; treat it as an option you hold, not a revenue line you forecast.
Where this argument is weakest
The strongest objection is not that the release theory is wrong. It is that it is out of date.
On this reading, 2026’s market is not a risk market at all. It is a supply market. Training rights appear in only about four in ten disclosed deals; the growth is in live retrieval and attribution; Amazon’s Times agreement buys summaries for Alexa; Wikimedia — with no claim to release — signs five enterprise partnerships in a single month; Microsoft’s marketplace meters usage rather than settling claims. Buyers, the objection runs, are purchasing freshness and reliability, and copyright exposure is a 2023 story that lawyers have not stopped telling.
That is largely correct about direction and it should change what most readers do, which is why supply is not a footnote here. But it does not displace the pricing claim, for four reasons. First, the money still follows claims: the largest agreements attach to parties who can litigate, and the Times licensed to Amazon while suing OpenAI. Second, supply contracts price like commodities, converging on the cost of the next-best supplier — which is why marketplace rates sit in pennies while bilateral deals with credible litigants sit in eight figures. Third, the two-to-ten-times certainty premium is a direct measurement of buyers paying for risk removal on top of supply. Fourth, the risk being transferred is broader than copyright: a buyer grounding a consumer assistant in a warranted feed is also purchasing a defence against being confidently wrong, which is a supply contract and a risk transfer in the same document.
Two developments would falsify the argument. If a marketplace with no litigation capacity behind it began clearing eight-figure aggregate payouts to long-tail publishers, the release component would be shown to be mispriced. And if a US appellate court affirmed training as fair use while bilateral deal values held steady, the cash would demonstrably be buying something other than risk. Neither has happened. Alsup’s ruling remains a single district-court decision that Anthropic chose to settle rather than appeal, so the fair-use question sits open above a $1.5 billion floor.
A worked deal: Fernshaw Vehicle Data
Fernshaw Vehicle Data is an invented but deliberately ordinary Coventry business: £4.4 million turnover, 31 staff, maintaining a UK used-vehicle residual-value dataset sold by subscription to dealer groups and motor-finance lenders. Its public asset is the Fernshaw Residuals Index, published free every month, which earns roughly 140 referring domains a year from trade press, brokers and consumer titles — the kind of recurring data-led asset that attracts citations without outreach.
In February 2026 an assistant developer building a car-buying agent approached them: £48,000 a year, perpetual, all purposes, no attribution, exclusive within automotive.
The No-Deal Test placed them in band three, not band four. The buyer could scrape the free index, and did; it could not obtain the underlying weekly transaction data anywhere else without building relationships with dealer groups, which was estimated at eighteen months. Their release was near-worthless — the index is published openly, and the Getty judgment made any offshore training claim unattractive. Their supply was the whole asset.
The displacement floor came to £26,400: 1,100 monthly answers naming Fernshaw, valued at £2 each. The original offer was below the floor before a single right was granted.
They restructured. Retrieval and display only, training carved out and separately priced at a figure the buyer declined; 24-month term with deletion on termination; named attribution with a live link on every surfaced answer, specified to the answer body rather than a source panel; no exclusivity; most-favoured-nation on rate; quarterly retrieval-and-attribution reporting. Signed in May 2026 at £91,000 a year.
Nine months on: cash up 90 per cent against the original offer, named appearances in the buyer’s product running at about 2,400 a month, and referring domains to the index up 21 per cent as trade titles wrote about the partnership. The honest negatives matter as much. Referral clicks from attributed answers were negligible — roughly 340 a month against 2,400 appearances, which is a brand outcome, not a traffic one. Legal fees consumed 14 per cent of first-year value. The buyer paused retrieval for seven weeks during a model migration with no contractual remedy, because Fernshaw had negotiated reporting but not a minimum-usage commitment. And the training carve-out cost them the deal with a second, larger buyer who would not proceed without it.
What to do on Monday
None of this requires a rights department. It requires deciding what you actually hold before anyone makes you an offer.
- Run the No-Deal Test on your three most-fetched assets and write the band number next to each. If everything lands in band one, stop preparing a media kit and start preparing an attribution ask.
- Separate your estate into release, supply, exclusivity and endorsement. Most sites discover they hold one of the four.
- Calculate the displacement floor for anything you would consider licensing. Twenty minutes, a citation tracker and a defensible presence value.
- Draft your attribution clause before you need it: surface, wording, link or plain text, answer body or source panel. Reuse it in every conversation.
- Audit any live agreement for perpetual grants and bundled training rights. Diarise renegotiation at the earliest break.
- Join a collective if it is free and non-exclusive, and forecast nothing from it.
- Keep funding the earned side. Third-party corroboration remains the input every buyer’s selection layer reads, whether or not a contract exists — the case for editorial placements, expert-source platforms and sponsorship-led coverage, paid insertions and launch-surface exposure is unchanged by any licensing outcome.
- Review how your material is structured for retrieval. A supply product needs stable schemas and predictable freshness, and the same discipline that helps you win concise answer placements makes a feed easier to license.
The uncomfortable finding underneath all of it is that quality is not what gets paid. Enforceability gets paid, and reliability gets paid, and both are orthogonal to how good the writing is. That is a poor reward system and it is the one in operation. The response is not to write worse; it is to stop expecting a licence to compensate for what the earned side is supposed to deliver. The fundamentals of why third parties point at you have not been repealed by any of this, and neither has the strategic case for building citable assets. Corroboration you did not pay for still governs whether you are selected, which is why what makes a product recommendable to a model matters more than any contract, and why the tooling you use to measure it should track named appearances rather than fees. The published data on how citations actually behave says the same thing every year.
Sell the release if you have one. Sell the supply if you can operate it. Ask for the credit either way, because it is the cheapest thing on the table and the only part of the deal that is still working in three years. And remember what the court actually decided: you cannot sell content to a company that already has it — you can only sell them the right to stop worrying about it, and the obligation to say your name.
Meta title: How to License Content to AI Companies: A 2027 Deal Playbook
| Meta description: A licence sells a release, not content. Use the Four Saleable Goods and the No-Deal Test to price an AI content deal and negotiate for attribution.
