Data-Scraping Litigation

Data-Scraping Litigation and Your Linkable Assets: A Risk Review

TL;DR

No plaintiff has yet won a web-scraping case by owning the data. The claims that survived 2024 to 2026 — Reddit’s anti-circumvention claim against Perplexity and SerpApi, Ryanair’s jury verdict against Booking.com — attached to something the publisher had built into its own servers before the crawl arrived: a technological measure, a login, a documented loss.

The claims that died asserted ownership. X Corp’s contract claims were held preempted by the Copyright Act; Meta’s terms did not reach a scraper that never logged in; the British Horseracing Board lost because it had invested in creating its data rather than obtaining it.

So a risk review of your linkable assets is not an inventory of what could be taken. It is an inventory of which claim elements you have already supplied, and which are still missing — plus the harder half, which is your exposure as a defendant when you are the one doing the collecting.

Nobody has won a scraping case by owning the data

Every scraping dispute is argued as a property dispute and none of them are decided as one. The plaintiff says: this is our data, on our servers, and they took it. The court then asks a series of questions that have almost nothing to do with ownership — was there a barrier, was there a promise, was there a measurable loss — and the case turns on the answers.

Run the record. In hiQ Labs v LinkedIn (31 F.4th 1180, 9th Cir. 2022), LinkedIn owned the servers, the traffic and the interface, and still lost the access argument: scraping a public profile is not access without authorization under the Computer Fraud and Abuse Act, the US anti-hacking statute. LinkedIn eventually prevailed on its User Agreement instead — which is the first clue. The instrument that worked was the promise, not the property.

In X Corp. v. Bright Data (N.D. Cal., 9 May 2024), Judge William Alsup dismissed X’s contract, trespass, misappropriation and unjust-enrichment claims. The contract claims were held preempted under section 301 of the Copyright Act, on the reasoning that enforcing terms of service over publicly posted content X does not own would do copyright’s work while shrinking what would otherwise be public. Alsup quoted the Ninth Circuit’s warning about the information monopolies that follow when a platform gets to decide who may collect data it does not own. The access claims failed separately, for want of any pleaded damage to X’s systems. Bright Data’s antitrust counterclaims — that X was monopolising the market in public-square data — then largely survived their own motion to dismiss, and the case ended in a stay on a settlement in principle. The scraper’s counterclaim outlived the platform’s claim.

In Meta Platforms v. Bright Data (N.D. Cal., 23 Jan 2024), Judge Edward Chen read Meta’s terms as governing your use of the products, held that a scraper collecting public pages while logged off was not a user at all, and noted that Meta had at some point after 2009 removed the clause binding anyone who merely accessed the site whether or not they were a registered member. The deletion was read as intent. Meta also argued that a former account holder stayed bound in perpetuity; the court disagreed.

In Europe, British Horseracing Board v William Hill (C-203/02, CJEU, 2004; Court of Appeal, [2005] RPC 13) killed the database-right claim of a body that had spent heavily on the fixtures data itself. Investment in creating data does not count towards the right; only investment in obtaining, verifying or presenting data that already exists does.

And the two live cases sharpen the point rather than blunting it. Ryanair won a unanimous CFAA jury verdict against Booking.com in Delaware in July 2024 — because a third party had gone through the password-protected myRyanair area — and then lost it in January 2025 when Judge William Bryson granted judgment as a matter of law: Ryanair had not proved the 5,000 US dollars of loss the statute requires. It is on appeal to the Third Circuit. Meanwhile, on 31 July 2026, Judge Paul Engelmayer let most of Reddit’s suit against Perplexity, SerpApi, Oxylabs and AWMProxy proceed, holding that Reddit could bring anti-circumvention claims under section 1201 of the Digital Millennium Copyright Act even though the copyrights in the posts belong to its users, because the statute lets any person injured sue.

The pattern is unmistakable. Every claim still standing attaches to an act of configuration. Every claim that failed attached to an act of authorship or an assertion of ownership. That is the review this article runs across your own assets — and if you build interactive calculators and data tools to earn links, it applies to you more sharply than to almost anyone, because those assets are designed to be extracted.

What is a scraping claim actually about? It is about the conditions under which a copy was made, not about who owns the thing copied. Each cause of action requires one fact that only the publisher can put in place, and only in advance: a barrier, an agreement, a record, or an investment of a particular type.

The creation penalty: your best asset is your weakest one

British and European publishers have one instrument Americans do not: the sui generis database right, a standalone right created by Directive 96/9/EC and enacted here by the Copyright and Rights in Databases Regulations 1997. It protects a database where there has been substantial investment in obtaining, verifying or presenting the contents — quantitatively or qualitatively, so analyst hours count as well as money. It lasts fifteen years from completion and renews on substantial change, which makes a continuously maintained dataset a rolling asset. The United States has no equivalent; Feist closed that door in 1991.

Then comes the trap. British Horseracing drew a line between obtaining data that already exists and creating data that did not exist before you made it, and put only the first inside the right. Football Dataco later confirmed the direction of travel. So the strength of your legal position runs in the opposite direction to the strength of your editorial position.

What counts as obtaining

  • Obtained: a compiled directory of public filings, a normalised table of other people’s published prices, a register assembled from scattered sources, a cleaned and verified aggregation of third-party statistics. The investment went into finding, checking and presenting things that already existed.
  • Created: an original survey you fielded, an index computed by your own model, the outputs of a calculator, a proprietary score, a benchmark generated from your own client data. However expensive, this investment made the data rather than gathered it.
  • Mixed, which is most real assets: a dataset where the raw rows were obtained and the derived columns were created. The right can subsist in the collection while the headline number you actually market has none.

The perverse consequence is worth stating plainly, because it inverts the usual advice. The original research study — the thing every link-building guide tells you to commission, and the thing that genuinely earns coverage — is the least protected asset class in Britain. The dull aggregation you would never brag about is the most protected. That is the creation penalty: the law rewards the collector, not the author, and it does so precisely because the right exists to subsidise assembly rather than invention.

Two things follow. First, if a dataset matters commercially, document the obtaining, verifying and presenting work separately from the modelling work, in timesheets and invoices, at the time. You cannot retro-fit creation into obtaining, but verification and presentation investment counts and is usually undocumented. Second, remember the geography: since 1 January 2021 a database made in the EEA does not attract the UK right and a UK-made database does not attract the EEA right, so a Manchester publisher and a Dublin publisher no longer hold the same instrument in each other’s market. That is now a live consideration in any international link building programme that publishes data across both.

Key takeaway

Investment in creating data earns no database right. Investment in obtaining, verifying and presenting it does. Your most original asset is your least defended one — so defend it with contract and configuration, not with a property claim you do not have.

This cuts against the commercial logic on purpose. The dataset that defends your rankings best is usually the one nobody else can replicate, and replication cost has nothing to do with legal protection. Both are true. They are simply different fights, and confusing them is how publishers end up paying for a claim they were never going to be able to bring.

The arming table: the element only you can supply

Each cause of action available to a scraped publisher has one element that no court, no defendant and no facts about the taking can supply. Only you can, and only before the event. A claim is therefore armed or unarmed as of today. You cannot install a lock after the burglary and then sue on it.

ClaimThe element only you can supplyWhere it livesArmed by default?
Circumvention (DMCA s.1201; CDPA s.296ZA)A technological measure that effectively controls accessEdge rules, bot management, rate limits, authNo — an open page has no measure, so no claim
Breach of contractAssent by the taker, not just published termsAccount signup, download gate, API key issueNo — a footer link binds almost nobody
CFAA / Computer Misuse ActA gate to exceed, plus quantified loss (5,000 US dollars)Login wall plus a costed incident logNo — public pages have no gate and nobody costs the response
Database right (UK and EU)Substantial investment in obtaining, verifying, presentingTimesheets, invoices, build recordsNo — and investment in creating the data does not count
CopyrightOwnership of the expression actually takenContributor and freelance agreementsPartly — you rarely own what your users or guests wrote

Read it as a budget rather than a menu. The cheapest row by an order of magnitude is the technological measure: a bot-management rule and a rate limit cost a morning of technical SEO work and convert an unprotected page into a page with a barrier. The most valuable row is assent, because a contract can reach uses that no property right touches. The most neglected is the loss figure, and Ryanair’s appeal exists entirely because of it.

Note what is not in the table: the quality of your content, the size of your audience, the egregiousness of the taking, and how obviously unfair it feels. None of those are elements of anything. This is why publishers who have been comprehensively strip-mined discover they have no claim, while a platform that lost the argument about ownership still has one.

The barrier is the claim

Section 1201 of the DMCA prohibits circumventing a technological measure — a technical control that gates entry — that effectively controls access to a copyrighted work. Britain’s equivalent, section 296ZA of the Copyright, Designs and Patents Act 1988, gives the person issuing copies to the public the same remedies as a copyright owner against someone who circumvents effective technological measures. Neither provision asks whether the taking was fair, transformative or commercially harmful. They ask whether there was a barrier and whether it was got around.

That is why Reddit’s case is the important one of 2026. Reddit does not own the posts; its users do, which is exactly why its earlier suit against Anthropic was pleaded in contract and unjust enrichment rather than copyright. Yet the anti-circumvention claim survived, because the alleged conduct was the evasion of Reddit’s and Google’s protections — rotating IP addresses and disguised requests at industrial scale — and the statute allows any person injured by a violation to bring an action. Standing followed the injury, not the byline. For a publisher whose content is assembled from many contributors, that is the most useful sentence in the current case law.

What counts as a technological measure?

What is a technological measure? It is a control that has to be defeated for access to occur: authentication, rate limiting, bot detection and challenge, token or signature checks, IP or ASN blocking, or a paywall. The test is functional. If a normal request succeeds and an automated one is stopped or slowed, you have a measure. If everything succeeds equally, you have a preference.

Does robots.txt count? No. A file asking well-behaved agents not to enter is a request, not a control, and ignoring it defeats nothing. This is the single most common misunderstanding in publisher circles: the crawl directives that shape how AI Overviews source your pages are a distribution setting, not a legal barrier. What changes the picture is a default block at the edge — Cloudflare’s 15 September 2026 default flip for the domains behind it, for example — because a default block is a measure, applied by you, in advance, and logged.

There is an honest limit. Section 1201 requires a copyrighted work behind the measure, and facts behind a lock are still facts. But the page is a work, the surrounding prose is a work, and the arrangement is at least arguably one, so the measure protects the container even where it cannot protect the numbers. It also does not require you to block anything you want indexed: a measure applied to an API feed or bulk export arms the claim while the HTML page stays wide open for the crawlers you actually want.

The preemption mirror: contract is strongest where your IP is weakest

Contract is the most powerful instrument available to a publisher, because it can reach uses no property right touches — resale, redistribution, training, competitive substitution. Its strength varies by jurisdiction, and it varies in the opposite direction to your intellectual property. Hold the two rulings side by side.

In America, weak IP weakens the contract. Alsup’s reasoning in X Corp. v. Bright Data was that a state-law contract claim doing the work of copyright, over material the plaintiff does not own, conflicts with the federal scheme and is preempted. The weaker your ownership of the underlying content, the more exposed your terms of service are.

In Europe, weak IP strengthens the contract. In Ryanair v PR Aviation (C-30/14, CJEU, 2015), the Court held that a database protected by neither copyright nor the sui generis right falls outside the Database Directive altogether — so the lawful-user protections in Articles 6(1), 8 and 15, which void contract terms that cut across them, simply do not apply. Ryanair’s flight data was unprotected, and that is exactly why its terms were enforceable against a screen-scraper.

The same unprotected dataset therefore supports a contract claim in London and invites a preemption ruling in San Francisco. If you publish from Britain and your likely counterparties are American, you are running two positions at once, and the drafting that helps in one forum is the drafting that draws fire in the other. Anyone advising on content licensing and pay-per-crawl terms should be sizing that split before writing a clause.

Assent is the scarce commodity

None of it matters without agreement, and this is where most publishers are unarmed. Meta turned on the fact that terms addressed to users do not bind visitors. A browsewrap — terms linked in a footer, accepted by nobody — is routinely unenforceable against an anonymous automated client. A clickwrap, where a specific act signifies acceptance, is routinely enforced.

On a public website there are only two events at which a stranger can be made to assent, and both are within your control:

  • An account. Registration produces a clickwrap and identifies the counterparty. It also carries a cost: gated content is content the answer engines cannot cite, which is the trade every publisher of interactive and scrollytelling assets has to price for themselves.
  • A download or key issue. The moment someone asks for the CSV, the dataset export or an API key, you can require an affirmative acceptance. This is the underused one. Put the licence on the file, not in the footer, and the asset that earns the most links becomes the asset with the strongest agreement attached.

Practical drafting follows from the two cases rather than from templates. Define the covered party to include automated and unauthenticated access, since Meta’s terms did not. Do not remove an access clause once you have one; the deletion was read against Meta and yours would be too. Version and date the terms, and log the acceptance event with a timestamp, so that assent is a record rather than an argument. State what happens on termination, because a former user’s obligations are not assumed to survive. And describe permitted extraction positively — what a lawful user may do — so that the prohibition has a shape a court can enforce.

Key takeaway

You cannot make a stranger agree to anything by publishing a page. You can make them agree at a download or a signup. Move the licence to the file and you convert your best link magnet into your best-documented contract.

The login line: you are more likely to be a defendant

The half of the risk review that nobody runs is the inbound half. Publishers building linkable assets are collectors. The rate table, the salary study, the pricing index, the state-of-the-market report and the competitor backlink analysis that seeds it all begin with acquisition, and the claim that reaches a small firm is a contract claim about how it collected, not a copyright claim about what it published.

One act changes the risk class, and it is not the volume of the crawl. It is the use of credentials. Logging in flips three switches at once:

  • Assent. You are now a user, and the terms address users. The Bright Data defence disappears the moment your session is authenticated.
  • A gate. Post-Van Buren, the American statute asks whether a gate was up or down. A login is a gate; a public page is not. Ryanair’s whole CFAA case existed because a partner’s agents went through the password-protected myRyanair area, and the jury found Booking.com vicariously liable for inducing it.
  • Identity. An authenticated session is tied to an account, so the counterparty’s logs identify you precisely. Logged-out collection is often anonymous in practice; logged-in collection never is.

The practical rule is unglamorous and holds up well: public, logged-out collection of published pages is the defensible posture and should be the default for anything that will be published. The moment a dataset requires a marketplace seller dashboard, a review platform’s back end, a client’s licensed tool export used beyond its seat terms, or a paid database’s bulk endpoint, you have changed the legal class of the asset and should know it before the analyst starts, not after the press release.

Worked example: the Pennine Freight Index

A Manchester logistics publisher builds a quarterly UK road-freight rate index as its flagship linkable asset. Two inputs. First, 240 published carrier tariff pages, crawled logged-out each week and normalised into a comparable table: 62 analyst hours over the quarter, about 4,100 pounds of documented obtaining, verifying and presenting work. Second, a 900-row extract of completed-load prices pulled from a marketplace’s seller dashboard using a client’s login, because the public pages do not show settled rates. The index model that turns both into the headline number cost about 11,000 pounds to build.

Now read the position. The 4,100 pounds buys a plausible database right in the compiled tariff table, because that investment went into obtaining and verifying data that already existed. The 11,000 pounds buys none, because the index is created. The 900 rows buy nothing at all — the publisher does not own them — and they carry the only real contract exposure in the project, because they were taken by an authenticated session behind a gate.

The timeline that follows is ordinary. The crawl starts on 14 January. The index publishes on 3 March and earns 71 referring domains in six weeks, which is the whole point of building it. On 22 April the marketplace’s counsel sends a cease-and-desist citing its terms and the authenticated access. On 6 May the publisher pulls the 900 rows, substitutes a smaller sample of publicly quoted rates, republishes with a methodology note, and preserves the collection logs. The index survives with a wider confidence interval. Nothing is filed.

Three lessons sit in that sequence. The cease-and-desist did the real work: it converted an ambiguous position into a documented one, and the correct response — stop, preserve, segregate, substitute — is what keeps a letter a letter. The 11,000-pound model is the asset the market values and the law ignores, which is the creation penalty in a single invoice. And the exposure travelled with the acquisition method, not with the published output, so no amount of care in the writing would have changed it.

One more warning from the same case law, for anyone tempted to send aggressive letters over public data. Bright Data’s antitrust counterclaims against X, alleging monopolisation of a market in public-square data, largely survived dismissal. Enforcement over material you do not own is not a free option; at scale it invites a competition theory in return, and that asymmetry is worth understanding before a specialist fires off a template.

When to sue nobody: extraction is often distribution

For a publisher whose business is earning citations, most extraction is the asset working. The engine that lifts three rows from your table and names you has done what the table was built for, and the sensible response to it is measurement, not litigation.

Citation versus substitution

The distinction that matters is citation versus substitution. Ask one question of every taking: does it create a competing destination? An answer engine that cites you converts extraction into a referral and a brand impression, which is what the zero-click traffic model is for measuring. A rival that republishes your compiled table as its own resource creates a substitute — a page that ranks against you, earns the links you would have earned, and is cited in your place. The first is distribution. The second is the case worth bringing.

European law encodes something close to that instinct. In CV-Online Latvia v Melons (C-762/19, CJEU, 2021), a specialised search engine indexed and reproduced content from a job-ads database. The Court accepted that this can amount to extraction and re-utilisation, but held that the database maker may prevent it where the activity adversely affects the investment in obtaining, verifying or presenting the contents, and directed national courts to weigh that harm against the interest in information services that create value for users. Harm to the investment, not the mere fact of copying, is the trigger.

There is one theory built for the modern pattern. Article 7(5) of the Directive, and regulation 16(2) of the 1997 Regulations, provide that repeated and systematic extraction of insubstantial parts may amount to extraction of a substantial part. An answer engine that takes three rows per response, ten thousand times, is the paradigm case — no single taking is substantial and the aggregate plainly is. But the theory is only as good as the evidence of the pattern, which means server logs, retained and analysed. Publishers who run a SERP-less audit of where their content already appears are, without meaning to, building the factual record that claim needs.

The rule of thumb

Licence the citers, sue the substitutes, ignore the crawlers. Suing an engine that cites you costs you the citation; suing a rival that replaces you costs you nothing you were earning. Most publishers get this backwards because the crawler is visible in the logs and the substitute is not.

The strongest objection to all of this

The best counter-argument is not that the law is unsettled. It is that none of this will ever be litigated by a business your size. Arming five claims is a cost with no realistic enforcement path; the claim that clearly survived in 2026 belongs to a company with a licensing programme and a litigation budget; and the technological measure that would arm your best row is usually deployed and logged by your CDN, not by you. On that reading, a risk review is theatre.

Most of that is right, and the honest concession is specific: for a publisher under a few hundred thousand pounds of revenue, at least three of the five rows will never be worth arming, and litigation is not the return on any of them. Anyone selling you a compliance project on the promise of a lawsuit is selling the wrong thing.

Where the return actually is

The return sits somewhere else, in three places.

  1. The letter. Nearly every one of these disputes ends before filing. The first question opposing counsel asks is not what was taken but what was in place — and a letter that names a technological measure, an accepted licence and a dated log gets a different reply from one that expresses disappointment.
  2. The price. An unarmed asset licences at zero, because there is nothing to settle. Every publisher deal signed in the last two years was priced against the cost and risk of taking the content anyway. Arming a claim is not preparation for court; it is the only thing that puts a number on the alternative.
  3. The defence. The exposure that actually arrives is inbound, and it is decided on records you either kept or did not: which sessions were authenticated, what the terms said on the day, when the letter came and what you did in the following fortnight.

And Ryanair remains the cautionary case for the optimists. A plaintiff can win the liability question unanimously and still lose the judgment because it never costed the response. Whatever else your team does, the incident log — engineer hours, mitigation spend, investigation time, dated — is the cheapest evidence in this entire field and the one nobody keeps. It is the same discipline that makes spam and anomaly detection work: the value is in having recorded the baseline before the event.

What to do on Monday

A morning of work covers most of the gap. Take your three most-linked assets and run them through this in order.

  1. List the assets and classify the investment. For each dataset, split the spend into obtaining, verifying and presenting on one side and creating on the other. The first column is your database right; the second is not. Do it in the timesheet, not from memory.
  2. Put a measure on the export. Rate-limit and bot-manage the CSV, the JSON endpoint and any bulk path, while leaving the HTML page open to the crawlers you want. That single change arms the cheapest row in the table.
  3. Move the licence to the file. Require an affirmative acceptance at download or key issue, timestamp it, and keep the accepted version. Terms in a footer are a preference; terms accepted at a download are an agreement.
  4. Fix the two clauses. Extend the covered party to automated and unauthenticated access, and say what survives termination. Do not delete an existing access clause.
  5. Start the loss log. A one-line template: date, incident, engineer hours, mitigation spend, investigation time. Six months of it is worth more than any opinion you could buy today.
  6. Audit the inbound side. List every dataset built with credentials. For each, find the terms, the account holder and the date. If an asset in production was collected behind a login, decide now whether you would defend it or replace the input.
  7. Separate citers from substitutes. Once a quarter, check where your compiled tables are appearing. Engines that cite you go on the licensing list; rivals republishing the whole table go on the enforcement list. Nothing else needs attention.
  8. Diarise the dates. Cloudflare’s default flip on 15 September 2026, the Third Circuit’s decision in Ryanair, and the next round of rulings in the Reddit litigation will each move one row of your table. None of them will move all five.

None of this is a substitute for advice from a solicitor on a specific dispute, and it is not offered as one. It is the configuration work that determines what that advice will be able to say when you need it — which, on the evidence of the last three years, is the part that decides the outcome. The rest of the discipline is unchanged: the fundamentals in our complete guide to link building, the tactics in our strategies guide, the tooling you use to monitor it, and the current benchmark data you measure it against. What has changed is that the assets themselves now have a legal configuration, and it is set in advance or not at all.

Two closing observations for the sceptical. Publishers who track link velocity and citation coverage already own most of the evidence base this field requires and have never framed it that way. And those investing in content credentials and provenance metadata are building the timestamped record that every one of these claims eventually turns on, whatever the reason they started. The same is true of anyone with clean render and crawl diagnostics or a working handle on how source diversity works in AI Mode: the instrumentation is already there. The gap is that nobody has pointed it at the question of what, exactly, they would be able to prove.

Leave a Reply

Your email address will not be published. Required fields are marked *

Agent Liability and Brand Risk Previous post Agent Liability and Brand Risk: Who’s Accountable When an Agent Errs
Governance Checklist Next post The 2027 AI-Era Link Building Governance Checklist