| Case | Jurisdiction | Core question | Status as of July 2026 |
|---|---|---|---|
| Bartz v. Anthropic | US, N.D. Cal. | Was training on books fair use, and does sourcing from piracy sites matter separately | Settled; $1.5 billion settlement given final approval July 20, 2026 |
| New York Times v. OpenAI and Microsoft | US, S.D.N.Y. | Does output memorization and training on paywalled news substitute for the original market | Active; discovery disputes over deleted logs, summary judgment briefing underway |
| Thomson Reuters v. Ross Intelligence | US, 3rd Circuit (on appeal from Delaware) | Does training a competing product on copyrighted headnotes qualify as fair use | On appeal; oral argument held June 11, 2026, decision pending |
| Getty Images v. Stability AI | UK, High Court | Does training abroad on scraped UK-hosted images infringe secondary copyright or trademark | Decided; judgment issued November 4, 2025, split verdict |
| Concord Music Group v. Anthropic | US, N.D. Cal. | Did Anthropic infringe song lyrics, including via alleged mass torrenting from pirate libraries | Active; amended complaint filed July 22, 2026, summary judgment briefing underway |
The Foundational US Ruling: Training Is Fair Use, Piracy Is Not
The single most important US ruling for builders remains Judge William Alsup’s June 2025 decision in Bartz v. Anthropic. It established a bifurcated framework that most subsequent cases have had to grapple with: training a large language model on legally acquired, copyrighted books can qualify as fair use under US law, because the transformation involved in learning statistical patterns from text is different in kind from redistributing that text. But the same ruling drew a hard line around how the training copies were obtained. Anthropic had downloaded millions of books from pirate libraries, including Library Genesis and Pirate Library Mirror, and Judge Alsup held that this piracy created liability entirely separate from any fair-use defense around the training itself.
That split — fair use may cover the training method, but does not launder the acquisition method — is the single most useful mental model for builders evaluating their own data pipelines in 2026. A dataset can be defensible on the “how you used it” question and still expose an organization to massive liability on the “how you got it” question.
Settlement Anatomy: What $1.5 Billion Actually Bought
On July 20, 2026, the US District Court for the Northern District of California granted final approval of the Bartz v. Anthropic settlement, described by Bloomberg as the largest copyright class action settlement in US history. The mechanics are worth understanding in detail, because they are likely to become a template for future AI copyright settlements.
| Metric | Figure |
|---|---|
| Total settlement value | $1.5 billion plus interest |
| Works covered by claims filed | 440,490 works, out of roughly 480,000 to 500,000 eligible |
| Approximate payout per author | About $3,100 per work, per class counsel estimates |
| Source of disputed copies | Library Genesis and Pirate Library Mirror, both pirate book repositories |
| Non-monetary requirement | Anthropic must destroy the original pirated files and any copies derived from them |
| Final approval date | July 20, 2026 |
Two things stand out for builders. First, the settlement is narrowly about acquisition, not about training methodology — Anthropic did not concede that training itself was unlawful, and the earlier summary judgment on fair use for training stood. Second, the destruction requirement sets a real precedent: courts are willing to order not just payment but the deletion of both original pirated files and downstream copies, which has implications for any organization whose training corpus includes material of uncertain provenance.
The Cases Still Being Litigated
New York Times v. OpenAI and Microsoft — the discovery war
This case has become less about the underlying fair-use question and more about what OpenAI’s own internal data reveals. In November 2025, a magistrate judge ordered OpenAI to produce a de-identified sample of 20 million ChatGPT conversation logs, an order affirmed by the presiding judge on January 5, 2026. By July 2026, the plaintiffs — a coalition led by the New York Times — asked the court to sanction OpenAI, alleging the company had misrepresented its ability to search training and output data and had deleted logs in violation of preservation orders. Summary judgment briefing closed April 2, 2026, but the discovery misconduct dispute means the substantive fair-use question has not yet been resolved by the court.
Thomson Reuters v. Ross Intelligence — the first appellate test
This case predates the generative AI boom in some ways — Ross Intelligence trained a legal research tool on Westlaw’s copyrighted headnotes to build a directly competing product, not a general-purpose chatbot. In February 2025, Judge Stephanos Bibas granted Thomson Reuters summary judgment against Ross’s fair-use defense, reasoning that a commercial, non-transformative use creating a direct market substitute is unlikely to be fair use even when the underlying copying is intermediate. Ross was granted an interlocutory appeal, and the Third Circuit heard oral argument on June 11, 2026 — the first time a US appellate court has directly reviewed a fair-use ruling in an AI training context. The outcome will matter well beyond legal search, because it tests whether “does the output compete with the original market” can override the more training-friendly reasoning seen in Bartz v. Anthropic.
Getty Images v. Stability AI — a split verdict from London
On November 4, 2025, the UK High Court issued the first UK judgment addressing generative AI training and copyright, in Getty Images v. Stability AI [2025] EWHC 2863 (Ch). Getty had abandoned its primary copyright infringement claims mid-trial, leaving a secondary infringement claim over Stability’s scraping of Getty images, which the court rejected — largely because the actual training occurred outside the UK, limiting the reach of UK copyright law. Getty did secure a narrow trademark win, with the court finding limited infringement under sections 10(1) and 10(2) of the Trade Marks Act for early Stable Diffusion outputs that reproduced Getty’s watermark, while dismissing the broader trademark dilution claim. The practical lesson for builders: cross-border training pipelines can meaningfully change which country’s copyright law even applies.
Concord Music Group v. Anthropic — the lyrics fight, now with a piracy angle
Music publishers including Concord Music Group, Universal Music Publishing Group, and ABKCO Music have pursued Anthropic since October 2023 over roughly 500 songs’ worth of lyrics. In January 2026, a related but separate suit was filed covering more than 20,000 songs and seeking over $3 billion, alleging mass torrenting of lyrics from pirate “shadow” libraries — echoing the piracy-specific liability theory that proved decisive in Bartz. A Second Amended Complaint was filed July 22, 2026, narrowing the claims to direct infringement and removal of copyright management information. Dispositive motions were scheduled to be fully briefed by June 8, 2026, and summary judgment practice is underway.
The emerging market-harm test
Across Bartz, the Ross Intelligence appeal, NYT v. OpenAI, and Getty v. Stability AI, courts are converging on a fact-specific question rather than a single bright-line rule: does the AI system’s output function as a substitute in the market the copyright holder controls. Training on legally acquired material for a genuinely transformative system leans toward fair use. Training that enables a direct competing product, or that relied on pirated source copies, leans away from it — regardless of how the training itself is characterized.
Outside the US: The EU’s Opt-Out Framework
The European Union took a fundamentally different legislative path years before most of the current litigation began. Article 4 of the 2019 Digital Single Market Directive created a general text and data mining exception, widely understood to cover AI training, but Article 4(3) gives rightsholders the ability to reserve their rights — an opt-out — which removes their works from the exception if they express that reservation in an appropriate, machine-readable way.
The EU AI Act built directly on top of that structure. Providers of general-purpose AI models are now required to maintain a policy for complying with EU copyright law, specifically including a process to identify and honor opt-outs from the text and data mining exception. Machine-readable protocols under discussion or in early adoption include robots.txt-style directives, manifest files, and metadata-based rights-reservation schemes, with the European Commission running a stakeholder process to converge on standardized machine-readable formats. Unlike the US litigation, which is resolving the question case by case through multi-year lawsuits, the EU has essentially legislated an answer — training is permitted unless a rightsholder actively opts out — and shifted the fight toward compliance mechanics and enforcement rather than whether an exception exists at all.
Japan’s Permissive Stance and Its 2026 Pressure Points
Japan remains the most AI-training-friendly major jurisdiction. Article 30-4 of Japan’s Copyright Act permits the use of copyrighted works for machine learning and data analysis purposes without the rightsholder’s permission, provided the use does not unduly harm the interests of the copyright owner, and that provision has not been narrowed as of mid-2026. What has changed is the pressure around the edges: Japan’s 2026 intellectual property strategic program keeps Article 30-4 intact for training itself, while adding transparency obligations, exploring creator compensation frameworks, tightening rules on web crawling, and considering new legislation to address voice imitation specifically. The Japan Copyright Agency’s updated guidance in 2026 held the training exemption steady while signaling that scrutiny is shifting toward what AI systems output and how they were deployed, rather than whether training itself was lawful.
| Jurisdiction | Legal mechanism | Opt-out available to rightsholders | Current posture |
|---|---|---|---|
| United States | Fair use doctrine, decided case by case | No statutory opt-out; litigation-driven | Training often defensible if lawfully acquired; piracy-sourced data carries separate liability |
| European Union | DSM Directive Article 4 text and data mining exception, refined by the AI Act | Yes, under Article 4(3), via machine-readable rights reservation | Legislated exception with a live opt-out compliance and enforcement process |
| Japan | Copyright Act Article 30-4 | Limited; use must not unduly harm rightsholder interests | Most permissive major regime, though output-side scrutiny is increasing |
| United Kingdom | Narrow existing TDM exception plus emerging case law | Effectively yes, since UK’s exception is narrower than the EU’s | Getty v. Stability AI narrowed the reach of UK claims over foreign training activity |
What Is Actually Settled Versus Still Being Litigated
For builders trying to make real data-sourcing decisions in 2026, it helps to separate what current case law actually resolves from what remains genuinely open.
- Settled: Training an AI model on legally acquired, legitimately licensed or purchased copyrighted text can qualify as fair use under the reasoning in Bartz v. Anthropic, at the district court level.
- Settled: Sourcing training data from piracy sites creates liability that exists independently of any fair-use defense around training methodology, and courts will order both payment and data destruction.
- Settled: The EU has a legislated text and data mining exception with an opt-out mechanism that GPAI providers must actively honor.
- Settled: Japan’s Article 30-4 training exemption remains intact, though output and deployment are drawing new scrutiny.
- Still litigated: Whether training that produces a direct commercial substitute for the copyrighted material — the Ross Intelligence fact pattern — defeats fair use even when the training itself is lawful.
- Still litigated: Whether output memorization of paywalled news content constitutes infringement distinct from the training process, the central NYT v. OpenAI question.
- Still litigated: How far a national court’s jurisdiction extends when training occurs abroad but scraping or output distribution touches the local market, per the unresolved edges of Getty v. Stability AI.
- Still litigated: Whether lyrics and other short-form creative text warrant a different fair-use analysis than long-form books, the live question in Concord Music Group v. Anthropic.
Common mistake
Assuming that because Bartz v. Anthropic found training itself could be fair use, any dataset is legally safe as long as the AI system does something transformative with it. The settlement and underlying rulings turn heavily on provenance — where the copies came from — not just on what the model does with them. Teams that treat “fair use covers training” as a blanket shield, without auditing how their corpus was actually acquired, are relying on only half the legal picture.
What worked
Organizations that fared best in 2026 built data provenance logs at ingestion time — recording exactly where each dataset or crawl came from, under what license or exception, and whether any rightsholder opt-out applied — rather than trying to reconstruct that history after a lawsuit was filed. Being able to show a clean acquisition trail, separate from the fair-use argument about training itself, has repeatedly been the difference between a defensible position and a piracy-style liability exposure.
Frequently Overlooked Details
- Acquisition and use are separate legal questionsFair use can cover how a model learns from data while acquisition of that data remains independently unlawful, as Bartz v. Anthropic made explicit.
- Destruction orders are a real remedyCourts are willing to order deletion of pirated originals and all derivative copies, not just monetary damages.
- Jurisdiction depends on where training happensGetty v. Stability AI turned partly on the fact that Stability’s actual training occurred outside the UK, limiting the reach of UK copyright claims.
- Market substitution is the emerging testRoss Intelligence lost its fair-use defense largely because its product directly competed with Westlaw’s headnotes, a stronger factor than the copying method itself.
- EU opt-outs require machine-readable signalsA rightsholder’s reservation under DSM Directive Article 4(3) only works if expressed in a format crawlers can actually detect and honor.
- Japan’s exemption is not unconditionalArticle 30-4 includes a limit against uses that unduly harm the rightsholder’s interests, which remains untested in major litigation.
- Discovery disputes can outlast the merits questionNYT v. OpenAI shows how allegations about deleted logs and misrepresented data access can stall resolution of the underlying copyright question for months.
- Settlements do not set binding precedentThe Bartz settlement resolves that specific case but does not bind other courts, so the Ross Intelligence appeal and NYT ruling could still diverge from its reasoning.
Glossary
- Fair use
- A US legal doctrine allowing limited use of copyrighted material without permission, evaluated across factors including purpose, nature of the work, amount used, and market effect.
- Text and data mining (TDM) exception
- A copyright exception, most developed under EU law, permitting automated analysis of copyrighted works, including for AI training, subject to conditions and possible opt-outs.
- Rights reservation (opt-out)
- A mechanism under EU law allowing copyright holders to exclude their works from the text and data mining exception by expressing that reservation in a machine-readable format.
- General-purpose AI (GPAI) model
- A category defined under the EU AI Act covering foundation models with broad applicability, subject to specific copyright compliance and transparency obligations.
- Class action settlement
- A negotiated resolution binding an entire defined class of claimants, such as the Bartz v. Anthropic settlement covering hundreds of thousands of books.
- Article 30-4
- A provision of Japan’s Copyright Act permitting the use of copyrighted works for data analysis and machine learning without permission, subject to a rightsholder-harm limitation.
Key Takeaways
- Bartz v. Anthropic established that training on lawfully acquired books can be fair use, while sourcing copies from piracy sites creates separate, serious liability.
- Anthropic’s $1.5 billion settlement received final court approval on July 20, 2026, covering more than 440,000 claimed works and requiring destruction of pirated files.
- NYT v. OpenAI has become dominated by discovery disputes over deleted logs and data access claims, delaying resolution of the core fair-use question.
- The Thomson Reuters v. Ross Intelligence appeal, argued June 11, 2026, is the first US appellate review of fair use in an AI training context.
- Getty Images v. Stability AI produced a split UK verdict: Stability won on copyright, Getty won a narrow trademark claim, and jurisdiction over foreign training mattered heavily.
- The EU has legislated a text and data mining exception with a rightsholder opt-out, shifting its fight toward compliance mechanics rather than whether training is permitted at all.
- Japan’s Article 30-4 remains the most permissive major framework for AI training, even as transparency and output-side scrutiny increase.
FAQs
Is training an AI model on copyrighted data legal in the United States?
It can be, under the fair use doctrine, if the training data was lawfully acquired. Bartz v. Anthropic found that training on legally obtained books qualified as fair use, but the same ruling held that training data pirated from sites like Library Genesis created separate, unrelated liability.
What was the outcome of the Bartz v. Anthropic settlement?
Anthropic agreed to pay $1.5 billion plus interest, covering more than 440,000 claimed works at roughly $3,100 per work, and to destroy the original pirated files along with any derivative copies. The US District Court for the Northern District of California granted final approval on July 20, 2026.
What is the New York Times v. OpenAI case actually about now?
While the case originally centered on whether training on paywalled news content and memorizing outputs constitutes infringement, by mid-2026 it has become dominated by discovery disputes, including allegations that OpenAI misrepresented its data access capabilities and deleted logs in violation of preservation orders.
Why does the Thomson Reuters v. Ross Intelligence appeal matter so much?
It is the first US appellate court review of a fair-use ruling in an AI training context. The district court found that training a directly competing legal research product on copyrighted headnotes was not fair use, and the Third Circuit’s decision, argued June 11, 2026, could set binding precedent well beyond legal research.
What happened in Getty Images v. Stability AI?
The UK High Court issued a split verdict on November 4, 2025: it rejected Getty’s secondary copyright infringement claim, partly because Stability’s actual model training occurred outside the UK, but found narrow trademark infringement in early Stable Diffusion outputs that reproduced Getty’s watermark.
How does the EU’s approach to AI training data differ from the US approach?
The EU legislated a text and data mining exception under the DSM Directive that generally permits AI training, but lets rightsholders opt out via a machine-readable rights reservation. The AI Act now requires general-purpose AI providers to actively identify and honor those opt-outs, shifting disputes toward compliance rather than case-by-case litigation.
Is it true that Japan allows AI companies to train on any copyrighted content?
Largely, yes, under Article 30-4 of Japan’s Copyright Act, which permits use of copyrighted works for machine learning without permission provided it does not unduly harm rightsholder interests. Japan’s 2026 policy updates keep this exemption intact while adding transparency and output-focused scrutiny.
What should builders actually do given all this ongoing litigation?
Maintain detailed provenance logs showing where every training dataset came from and under what legal basis, avoid any data source connected to piracy sites regardless of jurisdiction, and track opt-out signals in markets like the EU. Treat “fair use covers training” as necessary but not sufficient without a clean acquisition trail.
For related governance context, see our reports on digital rights management and protecting creator IP, EU AI Act compliance, and AI risk management frameworks compared. Builders assessing dataset transparency obligations may also find our piece on AI transparency reporting useful, alongside our companion articles on AI content provenance and detection limits and governing shadow AI tool adoption.
References
- Bartz v. Anthropic PBC, N.D. Cal. — summary judgment ruling, June 2025, and final settlement approval, July 20, 2026
- Copyright Alliance — “What to Know About the $1.5 Billion Bartz v. Anthropic Settlement”
- Locus Magazine — “Anthropic Settlement Receives Final Approval”
- The New York Times v. Microsoft and OpenAI, S.D.N.Y. — docket and discovery rulings, 2025 to 2026
- Variety — “New York Times and Other News Outlets Accuse OpenAI of Lying in Discovery”
- Thomson Reuters Centre GmbH v. Ross Intelligence Inc., D. Del. and 3rd Circuit — summary judgment opinion, February 2025, and appeal, Case No. 25-2153
- Getty Images (US) Inc and others v. Stability AI Limited, [2025] EWHC 2863 (Ch), UK High Court, November 4, 2025
- DLA Piper — “Getty Images v Stability AI: The UK High Court Decision Offers Guidance”
- Concord Music Group, Inc. v. Anthropic PBC, N.D. Cal. — docket and Second Amended Complaint, July 22, 2026
- Music Business Worldwide — “Music Publishers File Amended Lyrics Lawsuit Against Anthropic”
- Clifford Chance — “Copyright Compliance Under the EU AI Act for GPAI Model Providers”
- Global Law Experts — “Generative AI Copyright Japan (2026)”
