A USD 1.5 billion settlement has brought one of the most consequential copyright disputes in the generative AI industry to a close. Its significance lies in a distinction that every AI developer using copyrighted material must now confront: a court may treat the use of works for model training differently from the manner in which those works were acquired.
In Bartz v. Anthropic PBC1, authors and copyright owners alleged that Anthropic had downloaded millions of books from Library Genesis and Pirate Library Mirror and retained them in a central digital library. The dispute did not ultimately produce a single answer to whether the use of copyrighted books in artificial intelligence is lawful. Instead, the Court separated Anthropic’s conduct into distinct acts and subjected each to its own copyright analysis.
In June 2025, the United States District Court for the Northern District of California held that Anthropic’s use of books to train its large language models was fair use on the record before it. The Court also treated the one-to-one digitisation of lawfully purchased print books for an internal research library as fair use. It reached a different conclusion on Anthropic’s downloading and retention of pirated copies to build a permanent central library, holding that this activity was not protected by fair use at the summary judgment stage.
The surviving piracy-related claims proceeded toward class adjudication before the parties agreed to a USD 1.5 billion settlement. The final approval order described it as the largest copyright class action settlement in history and approved relief for owners of qualifying works included on a defined Works List2.
The litigation therefore presents a more precise question than whether AI training is fair use: can a transformative downstream use excuse the unlawful acquisition and retention of the copies used to achieve it?
Fair Use for Training Did Not Cure the Acquisition of Pirated Books
The Court’s analysis began by distinguishing the purpose for which Anthropic used the books from the way in which it obtained and retained them.
Anthropic had used copies of books in the process of training large language models. The Court regarded that training use as transformative because the books were not reproduced for readers in their original form but were used to develop systems capable of generating new text. On the particular evidentiary record, that use was held to be fair.
The Court separately considered books that Anthropic had purchased in print and converted into digital copies for storage and searchability in its internal library. Because the conversion was one-to-one and the physical copies were destroyed, the Court also treated that activity as fair use.
The position was materially different for books downloaded from LibGen and PiLiMi. Anthropic had acquired millions of pirated copies and retained them in a central library, making further copies of only some of them for model training while keeping the wider collection. The Court held that the later possibility of using a book for an otherwise fair training purpose did not justify the antecedent act of downloading and permanently retaining a pirated copy.
That separation is the litigation’s principal doctrinal contribution. Fair use is assessed by reference to the particular use under challenge. It does not operate as an enterprise-wide exemption capable of cleansing every act in the chain by which a copyrighted work enters an AI system.
An AI developer may therefore succeed in showing that model training is transformative while remaining exposed for the acquisition, copying or storage of the source material. The relevant legal questions arise at different points:
- Was the underlying copy lawfully acquired?
- Was an additional reproduction made?
- For what purpose was that reproduction used?
- Was the work retained for purposes beyond the allegedly transformative use?
- Did the company create a permanent library independent of its training activity?
The class proceedings and settlement followed this distinction. The settlement class was confined to beneficial or legal owners of the reproduction right in qualifying books contained on the Works List. The listed works were required to satisfy specified copyright-registration and identification criteria. Works outside that list were excluded, preserving claims concerning those works rather than releasing them through an uncertain class definition.
The final settlement created a USD 1.5 billion non-reversionary fund and was expected to yield approximately USD 3,000 per eligible work before deductions and allocation among rightsholders. Anthropic was also required to destroy the original files downloaded from LibGen and PiLiMi, together with copies originating from them, subject to legal-preservation obligations.
The scope of the release was equally deliberate. It covered claims relating to past piracy and copying of listed works up to the point of any AI output. It did not release claims based on past model outputs, future conduct on or after 25 August 2025, or works outside the Works List.
The settlement consequently resolved the liability risk associated with a defined historical dataset and defined past conduct. It did not provide a general judicial endorsement of Anthropic’s future training practices, nor did it determine the legality of AI-generated outputs.
For AI companies, the practical implication is that training-data governance must reflect the legal separateness of acquisition, retention, training and output generation. A company assessing only whether its final use is transformative may miss liability arising earlier in the data lifecycle.
That requires controls capable of establishing:
- the source from which each material dataset was obtained;
- whether the relevant copies were licensed, purchased or otherwise lawfully accessible;
- the purposes for which those copies were initially acquired;
- whether they were retained independently of model training;
- what derivative datasets or additional copies were created; and
- whether disputed material can be traced, segregated and deleted.
These are not merely compliance preferences. In Bartz, the source and retention of the books determined which claims survived even after the training use itself was held fair.
Conclusion
Bartz v. Anthropic should not be reduced to the proposition that AI training is fair use or that AI training infringes copyright. The litigation demonstrates why both formulations are too broad.
The Court treated model training, the digitisation of purchased books and the creation of a permanent library from pirated copies as legally distinct acts. Training was held fair on the record before the Court. The acquisition and retention of the pirated library were not insulated merely because some of those copies might later serve a transformative purpose.
The USD 1.5 billion settlement gives that distinction its commercial consequence. Even where the downstream use of copyrighted works is defensible, the route by which those works entered the developer’s systems may generate class-wide exposure on an exceptional scale.
For the generative AI industry, the central question is therefore no longer confined to whether training is transformative. It is whether each stage in the life of the training material-from acquisition and storage to model use and eventual deletion-can withstand independent copyright scrutiny.
The enduring lesson from Bartz is that fair use may defend a use. It does not necessarily defend the provenance of the copy on which that use depends.
Expositor(s): Adv. Aparna Shukla