TL;DR
Agent-error disputes are not being decided on fault. They are being decided on records — which party can produce what the machine said, to whom, how often, and at what cost.
The engine keeps a full record of its own conduct. The brand usually keeps a screenshot. Every 2026 outcome tracks that asymmetry.
Europe abandoned its fault directive and shipped an evidence regime instead: strict product liability with disclosure duties and presumptions, from 9 December 2026 — and it excludes the loss a brand actually suffers.
What follows: the four links of the exhibit chain, the arithmetic of proving that an answer recurs, and the capture protocol worth running in the first hour.
The answer everyone gives, and why it does not help you
Ask a law firm who is accountable when an AI agent errs and you will get a clear answer: the deployer. The organisation that switched the agent on cannot hide behind its autonomy. California’s AB 316, in force since 1 January 2026, removes the defence that an AI system acted of its own accord. Texas’s Responsible Artificial Intelligence Governance Act — a state statute covering AI developers and deployers — took effect the same day. Vendor terms push responsibility downstream to whoever put the system in front of a customer. The consensus is settled enough to be dull.
It is also, for most brands, an answer to a question they did not ask. Deployer liability describes what happens when your agent harms someone else. The exposure that keeps marketing directors awake runs the other way: a system you do not own, cannot audit and have no contract with tells a buyer something false about your company. There is no deployment for you to govern in that scenario. There is no agreement to sue on. You are not a party to anything — you are the subject of a sentence.
Between those two directions sits a third case that is quietly the most common of all: an engine repeats something you published, compresses it, and turns a true statement into a false one. You supplied the input. Somebody else supplied the error. Nobody has yet decided who owns the result.
Across all three, the question that actually decides outcomes is not who was at fault. It is who can prove what happened. Fault is an argument; the utterance is an exhibit, and in 2026 an exhibit is the scarce thing. This is not a rhetorical inversion — it is what the decided cases show, and it is what the one major piece of legislation to survive the period actually regulates.
What 2026 actually decided
What have courts decided about AI errors so far?
Very little about fault, and a surprising amount about who is speaking. Four decisions carry almost the entire weight of the field, and none of them turned on whether anybody meant to cause harm.
Munich: the answer is the engine’s own statement
On 28 May 2026 the Regional Court of Munich I (Landgericht München I, the civil court of first instance for the city) granted a preliminary injunction against Google in case 26 O 869/26, following an oral hearing on 23 April. Two Munich publishing houses had been linked by AI Overviews to scams, subscription traps and dubious business practices. The connections appeared in none of the cited sources. Google argued what search engines have successfully argued for two decades: that it displays third-party material automatically, does not adopt it, and is at most an intermediary.
The court refused the host privilege. An AI Overview, it held, is not a display of search results but Google’s own content, because summarising and connecting sources produces new and independent statements — some of which appear nowhere in the underlying material. The injunction carries a penalty of up to €250,000 per breach and Google was ordered to pay 80% of the costs. It is a first-instance order, Google says it is not final, and an appeal has been announced. It is also the only place in the world where a business has so far obtained relief against an answer engine over what that engine said about it — and it did so in a jurisdiction where a European market’s procedural rules hand claimants a fast prohibitory remedy that English law is structurally reluctant to give.
Georgia: a perfect transcript that lost on three grounds
In Walters v OpenAI (No. 23-A-04860-2), Judge Tracie Cason of the Gwinnett County Superior Court granted summary judgment to OpenAI on 19 May 2025. ChatGPT had told a journalist that the radio host Mark Walters was accused of embezzling from the Second Amendment Foundation. He was not, and no such case existed. The claim failed on three independent grounds: no defamatory meaning, because a reasonable reader in that journalist’s position could not have taken the output as an assertion of fact; no fault, neither negligence nor actual malice, with the court crediting OpenAI’s efforts to reduce hallucination; and no recoverable damages.
Read the second ground again. OpenAI won partly because it could document its own care. The claimant had the transcript; the defendant had the file.
Vancouver and Minnesota: the screenshot and the calendar
Moffatt v Air Canada (2024 BCCRT 149, 14 February 2024) remains the cleanest statement of deployer responsibility anywhere. The airline’s chatbot told a bereaved passenger he could apply for a bereavement fare retroactively; the policy page said otherwise. Air Canada argued, in the tribunal’s summary, that the chatbot was a separate legal entity responsible for its own actions. The tribunal rejected that, found the bot was simply part of the airline’s website, and awarded C$812.02. What is less often noted is that Air Canada led no evidence at all about the chatbot — how it worked, who built it, what it had been told. The passenger’s screenshot was the only account of the conversation in the room, so it became the facts of the case.
The largest claim in the field has yet to reach a merits ruling. LTL LED, LLC — trading as the Minnesota solar installer Wolf River Electric — sued Google after an AI Overview stated the company was being sued by the state attorney general for deceptive practices, a claim supported by none of the four sources it cited. Damages were put at $24.7m in the letter served with the complaint, then at roughly $110m to $210m in initial disclosures. On 9 January 2026 the case was remanded to state court because Google’s removal notice was filed late. Eighteen months of litigation, and the only thing decided was a deadline. Meanwhile Robby Starbuck’s claim against Meta ended in August 2025 with an arrangement to advise on the problem rather than a judgment, and his October 2025 complaint against Google in Delaware, seeking more than $15m, is pending.
Key takeaway
Four leading matters; zero rulings on whether an AI system was at fault for generating a falsehood. What each turned on instead: whose statement it was (Munich), what the reader could be shown to have understood and lost (Walters), whose record of the conversation existed (Moffatt), and a filing date (Wolf River).
The exhibit chain
Any claim arising from an agent’s error has to establish four things in sequence, and each one is a document rather than an argument. Call them the utterance, the occasion, the hand and the loss. The chain fails at its weakest link, and for a brand three of the four links are held by somebody else.
The utterance is the sentence itself, verbatim, in the frame it appeared in — with the disclaimers, the labels and the cited sources visible, because the frame is what decided meaning in Walters. The occasion is when it was generated, to whom, in what account state, and how often it recurs. The hand is which system produced it and on what build, since an engine that has shipped four model updates since March can honestly say the thing you recorded no longer exists. The loss is the money, traced to named counterparties who acted on the sentence.
| Link | What a claim needs | What a monitoring stack stores | Who holds the complete copy |
| The utterance | Verbatim text with citations, labels and disclaimers in frame | A mention flag, a sentiment score, sometimes a cropped screenshot | Nobody, unless you captured it |
| The occasion | Timestamp, audience, account and region state, recurrence rate | A weekly run of your own prompts from your own account | The engine, and the buyer who saw it |
| The hand | Engine, surface and model build at the moment of generation | The engine name, rarely the build | The engine |
| The loss | Named counterparties, dated decisions, quantified value | Nothing — this data never enters the tool | You, if anyone asked in time |
The exhibit chain: what a claim over an agent’s error requires, against what the brand-monitoring layer actually retains.
The distribution in the fourth column is the whole problem. On the loss link you hold the only copy and routinely fail to make it. On the other three the engine holds a better record than you do, and it holds it whether or not it ever discloses it. That asymmetry is not incidental to these cases. It is these cases — and it is why the strongest single move available to a brand is unglamorous: become the party with the better file.
The hand link deserves particular attention, because it decays fastest and nobody records it. Engines ship model and retrieval changes continuously and rarely version them publicly, so a claim about what an engine said in March is a claim about a system that no longer exists by June. The same volatility that makes any published figure in this field perishable makes an undated observation legally inert. Writing down the engine, the surface and the date costs seconds and is the difference between an anecdote and an exhibit.
Why a monitoring dashboard is not evidence
Brand monitoring in AI answers has matured fast, and the discipline of running a brand-safety monitoring system on a budget is now well understood, and the tool landscape has settled into a handful of credible monitors. What those systems produce, though, is an index of exposure. An index is not an exhibit. The gap between the two is where claims die.
Start with reproducibility. Thinking Machines Lab ran a single identical prompt a thousand times at temperature zero — the setting meant to eliminate randomness — and got eighty distinct completions. SparkToro repeated brand-recommendation prompts 60 to 100 times per platform and found that the same list of brands came back under 1% of the time. SE Ranking’s work on local queries in AI Mode found that only around 35% of domains recur between runs, with two-thirds vanishing. A 2026 arXiv study of citation variability across Perplexity, SearchGPT and Gemini put it formally: visibility figures are sample estimators drawn from a response distribution, not readings of a fixed value.
For measurement, this is a sampling problem with a known solution. For evidence, it is something worse. It means the most economical defence available to any engine — we cannot reproduce that — is usually true. It also means the answer your buyer saw was conditioned on their account history, their region and their session, none of which you can recreate. That is why the buyer who forwards you the screenshot is the most valuable witness in the file, and why multi-turn query chains are so hard to evidence: the damaging sentence often arrives at turn four, downstream of a conversation that belonged to someone else.
What is citation drift?
Citation drift is the tendency of an answer engine to cite different sources, and describe the same brand differently, on repeated runs of an identical prompt. It is a property of probabilistic generation and of the retrieval layer feeding it, not a fault. It has a legal consequence: a single observation proves that a system can say something, never that it does.
Two further gaps are worth naming because they are cheap to close. Monitoring tools capture the answer, not the frame — and the frame carried the day in Walters, where the model’s own warnings and the user’s scepticism defeated defamatory meaning before fault was reached. And almost no tool records what a deep research mode or an agentic browsing session did on the way to its conclusion, which is exactly the material that shows an error was generated rather than retrieved.
The reproduction floor: what it costs to prove an answer recurs
If a claim appears with probability p on any given run, the chance of missing it entirely across n independent runs is (1 − p) to the power n. Set that equal to 5% and you get the sample size at which absence starts to mean something: n ≥ ln(0.05) ÷ ln(1 − p). The numbers are less forgiving than most monitoring cadences assume.
THE REPRODUCTION FLOOR
Runs needed to be 95% confident of seeing a claim at least once, by its true frequency: 50% → 5 runs. 20% → 14. 10% → 29. 5% → 59. 2% → 149. 1% → 299.
The mirror image is the rule of three (Hanley and Lippman-Hand, JAMA, 1983): if an event does not occur in n observations, the upper bound of its 95% confidence interval is about 3 ÷ n.
So an engine that runs your query 20 times and reports no reproduction has excluded nothing above roughly a 15% rate. At 30 runs it has excluded nothing above 10%. Ask for the run count before accepting the finding.
Frequency is not a metric here. It is an element: an injunction restrains repetition, and repetition is a rate.
This matters beyond litigation. Under section 1(2) of the Defamation Act 2013 a body trading for profit must show serious financial loss, and loss scales with how many buyers encountered the statement — a function of frequency across the prompts your market actually asks. A claim surfacing on 4% of runs of one obscure phrasing and a claim surfacing on 60% of runs of the category question are the same falsehood and completely different commercial objects. Only one is worth spending money on, and no dashboard that reports a mention count will tell you which.
A worked example
Eastgate Compliance is a composite: a Sheffield fire-safety consultancy, 34 staff, £4.2m of annual revenue, roughly 70% of it from framework renewals with housing associations. On 3 September a prospect forwards a screenshot in which an assistant says Eastgate was “subject to an HSE enforcement notice in 2024”. It was not. The claim appears to have been assembled from a genuine notice against a similarly named contractor plus an unrelated trade-press piece.
The response takes four days. Eight prompt variants covering the questions procurement staff actually ask, run 60 times each across four engines from clean sessions in two regions: 480 observations, logged with full-page captures. The claim appears 11 times on one engine and once on another, and never on the remaining two. That yields an observed rate near 18% on the engine that matters, with an upper bound below 5% on the two clean ones — a distribution that supports a specific complaint about a specific surface rather than a vague grievance about AI. The loss trace runs in parallel: two framework renewals worth £96,000 paused, both in writing, both citing the enforcement claim. Total cost, at 2026 API prices and two analyst days, is under £900.
Eastgate never sues. It sends the engine a report with 480 observations behind it, corrects the two source pages that seeded the confusion, and gives its sales team a one-page rebuttal with the dates. The 18% rate is down to 2% within six weeks. The file cost less than a day of a solicitor’s time and did the work that no letter could have done without it.
The supply question: three positions, three records
Before deciding what to keep, decide which side of the error you are on. One question sorts it: what did you supply? The answer determines both your exposure and, more usefully, which document will be doing the work.
THE SUPPLY QUESTION
You supplied the deployment. Your assistant, your support bot, your feed served to an agent. You own the sentence, however it was generated. Your artefact is a retention log.
You supplied the input. Your published words were compressed by someone else’s engine into something untrue. Nobody owns the result yet. Your artefact is a dated, public, machine-readable record of what you actually said.
You supplied nothing. The claim was invented about you. You have no contract and no duty owed to you. Your artefact is a reproduction file, because every route that exists demands proof of repetition and loss.
Position one: you deployed it
Moffatt is the template and the deployer consensus has hardened around it. The practical work is retention, and it is thinner than most teams assume: the conversation log, the system prompt and its version history, the retrieval corpus as it stood that day, and the model build. Air Canada lost a case it might have argued because none of that was in the room. Two things sharpen the point in 2026. Vendor terms increasingly push responsibility to the deploying business, and insurers have started to leave: in January 2026 the Insurance Services Office, which publishes the standard forms most US commercial policies are built on, issued endorsements CG 40 47, CG 40 48 and CG 35 08, letting carriers exclude generative AI exposures outright, and W. R. Berkley has confirmed an absolute AI exclusion across directors’ and officers’, errors and omissions, and fiduciary lines. Coverage B, the line those forms remove first, is personal and advertising injury — the head that covers defamation. The same endorsement strips protection from the brand whose bot defames a competitor and from the brand defamed by one. There is a further trap in the European regime: a product that is substantially modified after supply brings its modifier back inside the manufacturer’s obligations, and a system that is retrained, re-prompted or re-grounded in production is modified more or less permanently. Deploying an assistant is not a purchase with a fixed date. It is a supply relationship you renew every time you change the prompt.
Position two: you supplied the input
This is the quiet majority of brand errors: a discontinued product still described as current, a price from an old landing page, a policy you changed. The engine is repeating you, badly. There is no defendant worth naming, and the fix is documentary rather than legal — which is why a deliberate AI training source strategy and clean, dated, machine-readable statements of fact do more here than any correspondence. Provenance work has the same shape: content credentials and the discipline of signed versus unsigned assets exist to make an assertion traceable to a moment and a signer, which is precisely what an exhibit is.
Position three: you supplied nothing
The pure fabrication — the invented enforcement notice, the imagined lawsuit, the merger that never happened. Here the hallucination-correction playbook is the practical route and the legal one is mostly theatre, with one exception: where the falsehood looks placed rather than generated, competitor misinformation in AI answers has a human author somewhere, and human authors have motives, records and assets. The distinction between an organic error and a seeded one is itself an evidential question, and it is answered by sampling, not intuition.
Does the new EU Product Liability Directive help a brand?
No, and the reason is instructive. Directive (EU) 2024/2853 — the revised Product Liability Directive, which member states must transpose by 9 December 2026 and which applies to products placed on the market after that date — makes software and AI systems products subject to no-fault liability, extends damage to destroyed or corrupted data and medically recognised psychological harm, and cannot be excluded by contract. But claimants are natural persons, and the recoverable heads are death, personal injury, private property damage and data loss. Pure economic loss to a business is outside it. A brand whose renewals are cancelled by a false answer has no claim under the one strict-liability instrument Europe passed.
What the Directive does regulate is proof. It obliges disclosure of technical evidence and creates rebuttable presumptions of defectiveness where a defendant fails to disclose or where technical complexity makes causation excessively difficult to establish. Set that beside the fate of its sibling: the AI Liability Directive, which would have harmonised fault-based claims, was listed for withdrawal in the Commission’s work programme on 11 February 2025 and formally withdrawn in the Official Journal on 6 October 2025, leaving fault to national tort law. Europe spent four years on agent liability and shipped an evidence regime after abandoning the fault regime. That is the clearest statement anyone has made about where these disputes are decided.
The British position, plainly
For a UK company the routes are narrow and every one of them is an evidence burden rather than a doctrinal one.
Defamation first. Section 1(1) of the Defamation Act 2013 requires serious harm to reputation; section 1(2) says that for a body trading for profit, harm is not serious unless it has caused or is likely to cause serious financial loss. In Lachaux v Independent Print Ltd the Supreme Court held that serious harm is determined by reference to the actual facts about the statement’s impact, not merely the meaning of the words. English law had already converted this tort into an evidential exercise years before answer engines existed. A company must prove loss, and must prove publication to identifiable readers in the jurisdiction — which is why a screenshot from your own test account, run from your own office, is close to worthless standing alone.
Malicious falsehood avoids the seriousness threshold but requires malice: knowledge of falsity or an improper motive. A stochastic system has neither, so the tort is close to unusable against an engine — though not against a competitor who prompted, screenshotted and circulated the output, a different defendant entirely.
Then the remedy problem. The relief the Munich publishers obtained — a prohibitory order before trial — is the one English law is least willing to grant: under the rule in Bonnard v Perryman, an interim injunction will not normally be granted in defamation where the defendant intends to justify the statement. So the fastest remedy in the field is the one a British claimant is least likely to get, and the slower routes all require the loss to be proven first.
One open question is worth flagging because it will be litigated. The single publication rule in section 8 of the 2013 Act assumes an identifiable first publication from which limitation runs. A generated answer has no publication date: it is composed on demand, for one reader, and may never be composed again. No English court has ruled on how limitation applies to a statement that is manufactured fresh on every occasion. Until one does, the safe assumption is the unhelpful one — that your evidence needs to be contemporaneous, because you cannot rely on the clock starting when you notice.
Key takeaway
Britain has no AI-specific liability statute; the clearest instance of Parliament allocating responsibility for autonomous decisions remains the Automated Vehicles Act 2024, and it did that for vehicles. Everything else runs through torts written for publishers and manufacturers, all of which ask a brand to produce readers, repetition and money.
Where this argument is weakest
The strongest objection is not that records are unhelpful. It is that they are not the constraint. Walters had a verbatim transcript, a named recipient and an unambiguous falsehood, and lost on all three elements anyway. Munich is a single first-instance injunction under appeal, in a jurisdiction with unusually protective personality rights and an unusually fast preliminary-relief procedure. On this reading the binding constraint is doctrine — meaning, fault, damage — and a brand with an immaculate file simply loses more expensively.
Half of that is right, and the half that is right is important: for most brands no claim will ever be worth bringing, and no filing discipline changes that. But look at what each of Walters’ three grounds actually turned on. Meaning was decided by the frame around the words: the model’s warnings that it could not open the link, its statement that the matter post-dated its knowledge cutoff, the journalist’s own scepticism. Fault was decided by OpenAI’s documented efforts to reduce hallucination. Damages failed because no reader could be shown to have believed it. Three doctrinal holdings; three evidentiary outcomes, and in the middle one the defendant won on the strength of its own file.
That is the asymmetry a brand can actually act on. You cannot change the law before your incident. You can change which party arrives with the better record — and the same file that would support a claim you never bring is what makes an engine’s correction channel respond, what an insurer needs before it argues about an exclusion, and what stops a sales team improvising a denial. Judged as litigation preparation, capture is usually a waste. Judged as the input to the three remedies that actually run, it is the cheapest work in the discipline.
The capture protocol
The first hour after a brand-damaging answer surfaces determines what is available for the next eighteen months. Almost nothing here requires a lawyer, and all of it degrades fast.
THE CAPTURE PROTOCOL
Capture the frame, not the sentence. Full-page image plus saved HTML, showing cited sources, AI labels and disclaimers. Cropped screenshots of the offending line destroyed the context that decided meaning in Walters.
Record the hand. Engine, surface, model build if exposed, date and time with timezone, region, and whether the session was logged in. Engines ship silently; the build you saw may not exist next month.
Preserve the third-party witness. The buyer’s original message is the scarcest item in the file, because it is the only proof of publication to someone who is not you. Ask for it the same day and keep the email intact.
Establish the rate before you complain. Fixed prompt set, fixed run count, both regions, misses recorded as carefully as hits — a documented zero is evidence too.
Timestamp the file externally. A published hash or any independent timestamping service costs nothing and answers the obvious question about when the capture was made.
On your own deployments, run the mirror image: conversation logs, system-prompt versions, retrieval corpus snapshots and model builds, retained on a defined schedule rather than a default one.
Two structural moves sit above the protocol. The first is contractual: if an agent speaks for you, the vendor agreement should oblige log access, retention for a stated period and reproduction on request, because the party that can reproduce the output controls the narrative about what it said. The second is corpus-level. Dated third-party publications — the coverage earned through reactive newsjacking, the quotes placed through journalist request platforms — are simultaneously retrieval inputs that can displace a false claim and independent records with a fixed date. That dual role is unusual and undersold. It is also the honest limit of the argument: a link is not a remedy, and treating link building strategy as reputational insurance is how brands end up disappointed. It is a corpus intervention with an evidentiary side effect.
The Monday checklist
- Write down which of the three supply positions each of your top five AI-visible claims sits in. The answer dictates whether your next hour goes into retention or reproduction.
- Fix the prompt set: eight to twelve questions your buyers actually ask, versioned, with the run count and regions written down before anyone looks at a result.
- Add frame capture to whatever monitoring you already run — full page and saved HTML, not a cropped image. This is a settings change, not a project.
- Create a one-page loss-trace template for sales: who paused, when, in their own words, and what it was worth. Most brands discover this data never existed.
- Ask every AI vendor you deploy one question: how long do you retain conversation logs, and can I get them? Put the answer in the renewal file.
- Check whether your liability policy carries a generative-AI exclusion. Renewals since January 2026 are where these endorsements appear.
- Audit the sources behind your worst recurring error before escalating anything — most position-two errors are your own stale pages, and correcting them is faster than any complaint. The mechanics of AI citation recovery and of measuring entity authority both start here.
None of this makes an engine accountable. It makes you the party who can say, precisely and with dates, what happened — which in a field with no fault standard, no settled defendant and no reliable remedy is the only durable position available. The tools change quarterly; what an engine chooses to recommend changes weekly; what AI browsers surface changes with each release. The file is the only artefact that appreciates.
Ask who is accountable when an agent errs and the honest 2026 answer is: whoever can prove what it said. That is almost never the brand — and it is the cheapest thing on this list to change. The reputational half of the work has always run on the same asset as the AI search visibility playbook and the defensive side of link building: a dated, external, retrievable record that somebody other than you can vouch for, which is, stripped of the jargon, what a link has always been. What is new is that a court now wants the same thing.
