TL;DR
No court anywhere has decided whether training a generative model on your articles is lawful. The first American appeal on the question was argued on 11 June 2026 and is still pending, and it concerns a legal research tool a judge called non-generative.
The decisions that have landed become consistent once you sort them by which layer of a page was taken. Facts moved freely. Expression did not. Whole files pulled from a pirate library produced the only cheque anyone has cashed.
That ordering is bad news for commercial publishers, because the layer an answer engine takes from an informational article is the layer copyright was built not to protect — and the remedy that made Anthropic pay $1.5 billion was gated behind a register almost no web publisher has ever joined.
The conclusion: plan as a supplier of access, not as a rights holder. Everything enforceable in 2026 sits in a contract, not in the Copyright Act.
1. The case in the headline has not decided anything
One sentence is doing an enormous amount of unearned work in content strategy decks this year: after the AI copyright rulings, publishers have leverage. Nobody has granted that leverage.
The New York Times sued OpenAI and Microsoft in December 2023. In April 2025 the Judicial Panel on Multidistrict Litigation pooled a dozen related suits into MDL No. 3143 — a multidistrict litigation, meaning one judge handles the shared pretrial work — before Judge Sidney Stein in the Southern District of New York, with Magistrate Judge Ona Wang on discovery. Judge Stein denied most of OpenAI’s motion to dismiss on 26 March 2025: direct and contributory infringement survived, the DMCA section 1202 claims were dismissed with leave to amend, and the argument that the 2019–2020 training claims were time-barred was rejected.
Everything since has been discovery. On 5 January 2026 Judge Stein affirmed Wang’s order compelling OpenAI to produce a complete twenty-million-conversation sample of ChatGPT logs rather than the subset OpenAI had proposed. The Times is asking for a permanent injunction and, under 17 U.S.C. section 503(b) — the provision allowing destruction of infringing articles — the destruction of models trained on its work.
What has not happened: any ruling on fair use, any finding of liability, any trial date. A motion to dismiss decides whether a complaint may be heard, not whether it is right. Reading it as a holding about AI training is like reading a planning application as a finished building.
What has NYT v OpenAI actually decided?
Nothing about training. It has produced pleading-stage rulings and a series of hard-fought discovery orders. The legality of using copyrighted text to train a model remains undecided in that case.
Nor is there an appellate answer waiting in the wings. The only United States appeal on AI training and fair use is Thomson Reuters Enterprise Centre GmbH v ROSS Intelligence (No. 25-2153), argued before the Third Circuit on 11 June 2026 in front of Judges Restrepo, Montgomery-Reeves and Bove. It is on interlocutory appeal from Judge Stephanos Bibas’s February 2025 decision that ROSS’s use of 2,243 Westlaw headnotes was not fair use, a decision Bibas expressly confined to a non-generative tool and himself described as hard under existing precedent. ROSS is leaning on the same circuit’s April 2026 fair use ruling in American Society for Testing and Materials v UpCodes. As of August 2026 the panel has not ruled.
Everything else is a district judge, in one district, on one record, about one defendant — persuasive at best, and binding on nobody else.
What the decided cases actually hold
| Case | Court, date | What it decided | What it does not settle |
| Bartz v Anthropic | N.D. Cal., Jun 2025; settled Jul 2026 | Training on lawfully bought books was fair use; building a library from pirated copies was not | Nothing about outputs. Settled, never appealed, so it binds no one |
| Kadrey v Meta | N.D. Cal., Jun 2025 | Fair use on this record; authors have no entitlement to a training-licence market | Thirteen authors, not a class. Market dilution untested |
| Thomson Reuters v ROSS | D. Del., Feb 2025; 3d Cir. pending | Fair use rejected for a non-generative legal research tool | No appellate holding yet; argued 11 Jun 2026 |
| NYT v Microsoft (CIR claim) | S.D.N.Y., Apr 2025 | Bullet-point abridgments of news articles were not substantially similar | One judge, one record — but the closest thing to law on summaries |
| In re OpenAI (outputs) | S.D.N.Y., Oct 2025 | Summaries of novels may be substantially similar enough to plead | A pleading standard, not a merits finding; fair use unaddressed |
| Getty Images v Stability AI | EWHC, Nov 2025 | Model weights are not an infringing copy; narrow trade mark win | Training legality untouched; the claims were dropped mid-trial |
| GEMA v OpenAI; GEMA v Suno | LG München I, Nov 2025; Aug 2026 | Memorisation is reproduction; the TDM exception does not cover training itself | German law, under appeal, and the claimant was a collecting society |
| Advance Local Media v Cohere | S.D.N.Y., 2025–26 | Publishers’ claims survived dismissal | A complaint that may be heard. Nothing more |
Amber rows are fact-bound first-instance decisions; red rows carry no precedential weight at all. There is no green row — nothing here has been affirmed on appeal anywhere.
2. Sort the rulings by which layer was taken and they stop contradicting each other
The commentary treats these decisions as a coin-flip — Alsup and Chhabria for the developers, Bibas and Munich for the rights holders, London somewhere in between. That reading survives only if you ignore what was actually taken in each case.
The cleanest demonstration comes from a single judge ruling twice. In April 2025, in N.Y. Times Co. v Microsoft Corp., 777 F. Supp. 3d 283, Judge Stein dismissed the Center for Investigative Reporting’s claim that Copilot’s bullet-point abridgments infringed its articles. The summaries were not substantially similar as a matter of law: they condensed non-copyrightable elements — the facts — and differed from the originals in style, tone, length and sentence structure.
On 27 October 2025 the same judge let output claims proceed in the consolidated authors’ case. ChatGPT’s summaries of A Game of Thrones plausibly crossed into protected expression because they reproduced the plot, the characters and the themes, and so carried the tone and feel of the original. Judge Stein distinguished his own news ruling in a footnote: those outputs had summarised facts. Not everyone is comfortable with the line: the copyright scholar Matthew Sag argued that if a 580-word summary can infringe a 694-page novel, the idea–expression distinction is in trouble.
Munich reached the same place from the opposite direction. On 11 November 2025 the Munich I Regional Court held in GEMA v OpenAI (42 O 14139/24) that memorising song lyrics in a model’s parameters and emitting them on request is reproduction, and that Article 4 of the EU’s DSM Directive — the text and data mining exception, which permits automated analysis of lawfully accessed works — covers the preparation of training data but not the training of the model itself. In August 2026 the same chamber extended the reasoning to a music generator in GEMA v Suno, asserting jurisdiction over training conducted in the United States, applying US law to it and rejecting fair use. Its memorisation test is the sentence to carry away: what matters is not how much of a work a model absorbs, but which part.
Every one of these decisions turns on which stratum of a document travelled, and the strata are not equally owned.
The four layers of a page (and who owns each)
Layer 1 — the facts. Prices, dates, counts, findings, the answer itself. Owned by nobody, by design: 17 U.S.C. section 102(b) excludes ideas, procedures and discoveries from copyright. This is the layer a retrieval system is built to extract.
Layer 2 — the selection and arrangement. Which forty-six suppliers, which nine metrics, in what order, under what taxonomy. Thin protection, but real: Westlaw’s headnotes were held copyrightable for the editorial judgement in them, and that is what sank ROSS. In the UK and EU this layer is also where sui generis database right lives.
Layer 3 — the expression. The sentences, the structure, the voice, the particular way the thing is put. Fully protected — and the only layer any court has actually enjoined: memorised lyrics in Munich, plot-and-character summaries in New York.
Layer 4 — the file. The article as an object, copied and retained. This produced the only money anyone has been paid: Anthropic’s library, not Anthropic’s answers.
Reading rule. You are paid for the layers you can prove were taken. Answer engines take layers 1 and 2 by design and touch layer 3 by accident. So copyright is strongest exactly where your traffic is not.
Run an informational article through that stack and the result is uncomfortable. A buyer’s guide, a rate benchmark, a how-to, a comparison table — the entire value of these is layer 1 with a little layer 2 holding it together. Which is precisely why they get summarised, and precisely why the summary is lawful. The same property that makes a page useful enough to be cited in AI Overviews and the answer surfaces around them makes it a bundle of unownable facts.
There is a second inversion buried in the Munich reasoning. The court, following the computer-science evidence before it, linked memorisation to how often a work appears in the training set: popular songs are absorbed because they are everywhere. The only infringement theory that has actually won anywhere therefore depends on redundancy. Genuinely original, single-source material — the sort a serious content programme exists to produce — is the least likely to be memorised and so the least likely to generate a claim. Copyright, on the current evidence, protects the duplicated and abandons the novel.
Key takeaway
The question is never “was my content used?” It is “which layer moved?” Facts moving is not a wrong the Copyright Act recognises, however much revenue moves with them. Before you buy an opinion, locate your grievance on the stack — most commercial publishers discover theirs sits on layer 1, where there was never a right to lose.
3. Where something is owned, the remedy is still gated
Assume you clear the layer test: a competitor’s model emits your prose verbatim. You now meet the machinery that decides whether owning a right is worth anything, and it is administrative rather than moral.
The Anthropic settlement is the clearest specimen because its eligibility rules are public. On 20 July 2026 Judge Araceli Martínez-Olguín of the Northern District of California granted final approval to the $1.5 billion settlement in Bartz v Anthropic — the largest copyright class settlement on record, with a 91.3% claims rate and fifty-three objections overruled. Anthropic must destroy the files it took from LibGen and the Pirate Library Mirror. The release runs only to conduct before 25 August 2025 and expressly does not cover claims about outputs.
Now the part nobody quotes. The class was not “authors whose books were used.” It was the owners of works on a Works List of 482,460 titles, each of which had an ISBN or an ASIN and a United States copyright registration obtained within five years of first publication. Two administrative facts, and neither of them is authorship.
Why registration is the whole mechanism
Three provisions do the work. Section 411(a) makes registration a precondition to suing on a US work, and Fourth Estate v Wall-Street.com (2019) confirmed that registration means the Copyright Office has acted on the application, not merely that you filed it. Section 412 then bars statutory damages and attorney’s fees entirely unless registration preceded the infringement or fell within three months of first publication. Section 504(c) supplies the numbers that make litigation viable: $750 to $30,000 per work, rising to $150,000 where infringement is wilful. And section 410(c) explains the five-year condition on that Works List — a certificate obtained within five years of publication is prima facie evidence that the copyright is valid, which is what makes ownership provable at industrial scale.
String them together and the settlement stops looking like a valuation of literature. 482,460 works at the statutory minimum of $750 is roughly $362 million of exposure before anyone argues about wilfulness; the court noted that the per-work payment of about $3,000 is four times that minimum. The headline number was manufactured by multiplying a floor by a countable set of registered works. Remove the register and the multiplication has nothing to work on: you are left with actual damages, and the provable actual loss on a single unregistered blog post is indistinguishable from zero.
British publishers face a different gate to the same room. The UK has no registration system, so there is nothing to be late for — but there are also no statutory damages, only actual loss, an account of profits, or additional damages under section 97(2) of the Copyright, Designs and Patents Act 1988 in flagrant cases. Under the Berne Convention a UK work can be sued on in an American court without a US registration; section 412 still withholds statutory damages and fees. A British claim can get through the door and not to the money.
And the British evidential gate is the one Getty walked into. In Getty Images (US) Inc v Stability AI [2025] EWHC 2863 (Ch), Getty abandoned its primary copyright and database right claims mid-trial because it could not establish that training had happened in the UK, leaving only secondary infringement — which failed when Mrs Justice Smith held that Stable Diffusion’s weights never contained or stored the images, so the model was not an infringing copy. Rights survived; proof did not.
The government then declined to fix the proof problem. On 18 March 2026, meeting its deadline under sections 135 and 136 of the Data (Use and Access) Act 2025, DSIT, the Intellectual Property Office and DCMS published the Report on Copyright and Artificial Intelligence and its economic impact assessment. After more than 11,000 consultation responses, the answer was: no new legislation, no new regulator, existing law applied by existing courts. The December 2024 preference for a text and data mining exception with an opt-out was abandoned; ministers told the House of Lords in January 2026 that expressing that preference had been a mistake. On transparency, technical standards, licensing and enforcement the Report largely undertakes to monitor and to work with industry. For a UK site owner, that is a decision: the licensing and pay-per-crawl decisions you make on your own domain are the only compliance surface you control.
The two prices — run both before you spend anything
Price A, the litigation floor. F = n × s, where n is the number of your works you can prove you own and that were registered in time, and s is the statutory minimum of $750. F rises with volume. For almost every commercial content site n is zero, so F is zero no matter how much was taken.
Price B, the access rate. r = V ÷ N, where V is a licence’s annual value and N is the size of the corpus it buys. r falls as N rises, because the buyer is paying for coverage, not for craft.
The reading. Volume is an asset in court and a liability at the negotiating table. You cannot reach Price A without a register; you cannot reach a meaningful Price B without a corpus larger than your own. If F is zero and r multiplied by your own output is less than the cost of the meeting, you do not have a copyright strategy. You have an access strategy, and you should stop paying lawyers to look for the other one.
4. The licensing market prices corpora, not articles
The standard advice — get your content ready to license — collides with the second ruling from June 2025. In Kadrey v Meta, Judge Vince Chhabria granted Meta summary judgment on fair use and, along the way, held that the authors were not entitled to the market for licensing their works as AI training data. He called both of their theories losers: the model could not reproduce enough of their books to matter, and a lost training licence was not a market copyright protects. The theory he thought could win — dilution, the flooding of a market with machine-made substitutes — they barely pursued.
So the market people are told to prepare for has, in the one case that examined it, no protected status. It exists anyway, and its prices are public. OpenAI’s reported agreement with News Corp is worth up to $250 million over five years, in cash and technology credits, covering the Wall Street Journal, Barron’s, MarketWatch, Investor’s Business Daily, the New York Post, The Times, The Sunday Times, The Sun, The Australian and dozens of Australian regional titles. News Corp added Meta in March 2026 at up to $50 million a year for three years. Axel Springer is reported at roughly $13 million a year and Amazon’s deal with the New York Times at $20–25 million a year — close to one per cent of the paper’s revenue. Reddit disclosed $203 million of aggregate data-licensing contract value in its IPO filing.
Do the division. Assume, generously, that News Corp’s licensed titles publish 250,000 items a year between them: the largest content licence ever signed is worth about $200 per item per year. Divide Reddit’s disclosed contract value by a corpus counted in billions of comments and the per-item price rounds to nothing. These are not prices for works. They are prices for coverage.
Key takeaway
Statutory damages multiply by your volume; an access licence divides by it. A 5,000-article publisher is therefore too small to sue profitably and too small to sell directly — which is why the only entity to win an injunction, damages and a disclosure order against OpenAI so far was GEMA, a collecting society acting for thousands of writers.
That is the structural conclusion, and Britain already has the machinery: the Copyright Licensing Agency licenses text and image reuse on behalf of the Authors’ Licensing and Collecting Society, Publishers’ Licensing Services and DACS, and ALCS alone has distributed over £750 million to more than 130,000 members. The mandate those bodies hold is for reuse by businesses and institutions, not for AI training, and whether it is extended is the live question — the March 2026 Report floated a creative content exchange. For a publisher below roughly twenty thousand items, membership of something larger is the only realistic route into a licence.
5. What the docket is actually good for is evidence
Three years of litigation has produced very little law and a remarkable amount of fact, and the facts are free. Twenty million ChatGPT conversations are now in a court file. A Works List showed exactly which 482,460 books had provable ownership. An English judge examined a diffusion model and found no copies inside the weights. Munich ordered OpenAI to disclose the scope of its use of the works and the revenue earned from them. Treat the docket as a data source rather than a rights regime and it starts paying for itself.
The most valuable thing in it is a roadmap nobody is following. Judge Chhabria told the next plaintiff precisely what to bring, and market dilution is the one theory that does not require your expression to have been copied at all — which makes it the only theory available to publishers whose grievance sits on layer 1. It requires evidence that only you can create, and only if you create it before you need it.
The dilution file — four fields, one folder, opened this quarter
1. The market. Name the thing you sell and to whom. Not “traffic” — a subscription, a report, a dataset seat, a sponsorship. A market has customers and a price.
2. The price. Your rate card, dated, with the history. Dilution is a claim about a price being pushed down, which is impossible to argue without one.
3. The substitute. The specific machine-made thing a buyer took instead: prompt, date, engine, output, captured and stored. Screenshots age well; memories do not.
4. The displacement. A dated series from your own logs and your own renewals across the same window. Not a vendor index — your numbers, on your customers.
The industry-level backdrop is well documented and will not carry your claim. Chartbeat data published on 17 March 2026 via Axios found that small publishers — 1,000 to 10,000 daily page views — lost 60% of their search referrals over two years, mid-sized publishers 47%, and that Google Search referrals across more than 2,500 news sites fell 34% between December 2024 and December 2025; AI chatbot referrals grew more than 200% and still account for under 1% of publisher referrals. SparkToro’s June 2026 analysis of Similarweb clickstream data put the zero-click share of Google searches at 68% for January to April 2026, against 60% in 2024.
Every one of those figures is somebody else’s aggregate. They establish that a trend exists, which no defendant will contest; they say nothing about whether your price moved. That is why the zero-click traffic model you build from your own analytics and a SERP-less audit of your own demand are worth more than any published index — and why the published link building statistics belong in your introduction, not your evidence. Whatever tooling you track this with should be pointed at your own renewals first.
6. A worked example: what the stack costs in practice
Coldbrook Materials Review is a composite, with the figures made explicit so the arithmetic can be checked. Sheffield-based, founded 2013, £2.1m revenue — £1.35m subscriptions, £750k sponsorship — with 4,900 published articles and a monthly Ready-Mix Index compiled from 46 suppliers.
September 2025. An editor finds an assistant returning that month’s index values, unattributed and unlinked. Separately, two paragraphs of a 2019 technical explainer come back close to verbatim. November 2025. A City firm quotes £68,000 for an opinion and a letter before action. January 2026. Instead, the editor runs the layer test. The index values are layer 1, and the CIR ruling says a condensation of facts is not substantially similar. The choice of 46 suppliers and the construction of the series is layer 2 — arguable, possibly a database right given the investment in obtaining and verifying it, and immediately into Getty’s problem of proving where the acts occurred. The two verbatim paragraphs are layer 3 and the only viable claim in the building.
February 2026. The two prices. Registered works: zero, so F = 0 and section 412 removes statutory damages and fees from the one good claim. Actual loss on a 2019 explainer: unquantifiable. At $200 an item — the rate implied by the largest deal ever signed — 4,900 items would be worth under $1m a year in theory and nothing in practice, because no buyer negotiates for a 4,900-item corpus. The £68,000 was the only certain number on the page.
March to July 2026. What they did instead. The index moved behind an authenticated endpoint with a per-seat licence, published machine-readable terms and access controls enforced at the technical level; the explainers stayed open and free. They opened a dilution file. They registered the twelve monthly index reports in the United States within the three-month window. By July: 41 seats at £3,600 a year, £147,600 of licence revenue that did not exist in January.
The honest cost. The index landing page’s AI citations fell to approximately zero and its organic sessions dropped 62%, because a gated page cannot be quoted. Two sponsors queried the paywall and one renewal was lost. They converted an unenforceable claim into a priced contract and paid for it in visibility — which is the trade, and not one every publisher should make. The explainers were left open for exactly that reason: they are the quotable, snippet-shaped material that keeps the brand in answers, and they were never protectable anyway.
7. The strongest objection to all of this
Here is the best version of the counter-argument. The money is already moving. Anthropic paid $1.5 billion. OpenAI has signed something like two dozen publisher agreements covering 160-plus outlets. GEMA has won twice in Munich. Whatever the doctrine says, the practical direction of travel is that content gets paid for, and a publisher should be positioning to be paid rather than writing copyright off as an administrative dead end.
Conceded without qualification: the threat of litigation, not the outcome of any of it, created the licensing market. Three answers survive it.
- The payments prove the eligibility point. Every cheque went to a counterparty that could deliver an enumerable corpus and an indemnity — a Works List, or a masthead estate. That is evidence for the argument above, not against it.
- The biggest settlement in the history of copyright carved out the relevant question. Anthropic’s release covers conduct up to 25 August 2025 and expressly excludes claims about outputs. What an engine may do when it answers using your page was not resolved; it was reserved.
- Nobody knows the direction of travel. The only court positioned to make binding law heard argument on 11 June 2026 and has not ruled, and the decision under review concerns a non-generative tool.
What would change the analysis
Two things, and naming them is the honest way to hold a position. If the Third Circuit affirms Judge Bibas on a functional-substitute rationale — a product that answers the question the original answered harms the market for it — then factor four acquires a shape that fits informational publishing better than it fits novels, and the layer 1 conclusion above weakens considerably. And if any court accepts market dilution on a properly built record, layer 1 stops being a dead end, because dilution is not a claim about copied expression. Both are live. Neither has happened. Anything built on either is a bet and should be sized like one.
8. Two honest negatives
Registration is cheap, and I have just spent a section explaining why it usually will not save you. That needs narrowing. Register anything that can be copied verbatim: datasets, photography, illustration, paid reports, original calculators and interactive assets. Do it inside the three-month window, because section 412 is a timing rule and not a filing rule, and you gain the section 410(c) presumption. My claim is narrower than “do not bother” — registration cannot create a right over being summarised, because summarising facts was never an infringement to begin with.
Access control has a measured cost, and it lands on discovery. Gating is the lever that actually works, and it removes you from the surfaces where citations now happen. The Coldbrook numbers show the shape of the bill: licence revenue up, citations on the gated asset gone. Decide asset by asset — some pages exist to be quoted, some exist to be bought, and the mistake is treating a whole domain as one or the other. This is not legal advice, and no substitute for an adviser who knows your jurisdiction and your contracts.
9. What this changes on Monday
- Run the four layers over your three most valuable pages. Write down which layer your grievance sits on. If it lands on layer 1, stop — there was never a right there to lose.
- Compute F. How many of your works carry a registration that survives section 412? For most sites the answer is zero and confirming it takes ten minutes.
- Register what is genuinely copyable, on publication. Datasets, reports, imagery and original assets — inside three months, not when you notice a problem.
- Publish machine-readable terms and log crawler behaviour. A future claim needs an act and a date; how crawlers actually render and fetch your pages determines what your logs can prove.
- Open the dilution file. Rate card, dated prompt captures, your own log series and renewals. Four fields, one folder, this quarter.
- Price access, not rights. A per-seat licence on the one asset nobody can restate beats an opinion letter on the whole archive — and a structured feed or API is how you sell it without leaking it.
- If your corpus is under roughly 20,000 items, stop budgeting for a direct licence. Budget for membership of something larger, and ask what mandate it actually holds.
None of this replaces the work that earns the citation. The strategies that build authority, the fundamentals underneath them and what a backlink actually is still decide whether you are in the retrieval set, and none of that changed because a judge in Manhattan is reading chat logs. How AI Mode spreads its citations and how deep research modes assemble a source list will move your visibility this quarter; a ruling will not.
Provenance work does the one thing copyright cannot. Content credentials, C2PA signing, the labelling rules attaching to AI output and the difference between signed and unsigned assets create a record of who published what, when — and evidence, not entitlement, is the scarce input in every case above. It is also the cheapest thing on this list, which is why the premium on demonstrably first-hand material keeps rising while the litigation stalls.
The American question is whether your words were copied. The British question is whether you can prove where it happened. The engine asks neither. It takes the facts, which were never yours, and the only reply the law leaves you is a price.
