Why Are Some AI Companies Buying Millions Of Books...And Then Destroying Them? The AI Race For Books.
- 6 hours ago
- 10 min read
Anthropic, who owns Claude AI, physically cut apart millions of purchased books while building a private digital library and other major AI companies are tied primarily to digital collections obtained from shadow libraries, research datasets or earlier scanning programs.
Court records confirm that at least one major artificial-intelligence company purchased millions of physical books, removed their bindings, cut apart their pages, scanned them and discarded the paper originals. That company was Anthropic, developer of the Claude AI models.
The same records show that Anthropic had previously downloaded more than seven million digital book files from sources the court described as pirate libraries. Its later physical purchases were part of an effort to obtain books through a different route. The evidence involving Meta, OpenAI, Microsoft, NVIDIA and Google is different. Court orders and company disclosures connect those firms to digital collections including Library Genesis, Anna’s Archive and Books3. Google also operated a documented program that scanned millions of physical library books, but the public records reviewed for this article do not establish that those companies ran an Anthropic-style operation in which purchased books were cut apart and discarded.
There are two related stories: the destruction of physical books to create private digital copies, and the use of enormous digital libraries containing works that were allegedly or demonstrably obtained without authorization.
Anthropic First Obtained Millions of Digital Books
The clearest account appears in Bartz v. Anthropic, a copyright case filed in the U.S. District Court for the Northern District of California. According to Judge William Alsup’s June 23, 2025 order, Anthropic co-founder Ben Mann downloaded Books3 in early 2021. The collection contained 196,640 books assembled from unauthorized copies of copyrighted works. In June 2021, Mann downloaded at least five million books from Library Genesis, commonly called LibGen. In July 2022, Anthropic downloaded at least two million more from Pirate Library Mirror, or PiLiMi.
The judge concluded that Anthropic had obtained more than seven million pirated book copies. These were full text books stored in formats including PDF, EPUB and plain text, not merely bibliographic information, summaries or excerpts. Anthropic retained the collections in a centralized research library and selected different sets for different research and model development purposes. The court said Anthropic also kept books that might never be used in training, indicating that the company was building a permanent corporate collection that could be searched and reused for future projects.
Anthropic Later Began Buying Physical Books
By 2024, Anthropic was concerned about the legal risks associated with pirate libraries. In February of that year, the company hired Tom Turvey, formerly the head of partnerships for Google’s book-scanning project, and tasked him with helping Anthropic obtain a broad collection of books. Turvey made limited contact with publishers about possible licensing arrangements, according to the court, but those conversations did not produce the collection Anthropic sought. His team instead contacted major book distributors and retailers about purchasing print copies in bulk.
Anthropic then spent what the court described as “many millions of dollars” purchasing millions of physical books, frequently used copies. The books were sent to service providers that removed their bindings, cut the pages to dimensions suitable for production scanners, scanned the pages and used optical character recognition to create machine-readable text. The paper pages and remaining binding materials were then discarded.
Each purchased book was replaced by an internal PDF containing page images and searchable text. Anthropic also created bibliographic records so the books could be cataloged and located inside its research library. The physical books were therefore not destroyed after Anthropic finished training an AI model on them. They were destroyed during the process of creating the digital scans, which the company retained whether or not a particular book was eventually selected for training.
The Operation Was Known as Project Panama
Additional details emerged through thousands of pages of exhibits unsealed in the litigation. The physical book acquisition and scanning effort was internally known as Project Panama. A vendor proposal described the possible conversion of between 500,000 and two million books during a six month period. The total number ultimately scanned and the full cost remained partly redacted, but the court independently established that Anthropic purchased and destructively scanned millions of books.
The documents describe an industrial process rather than occasional office scanning. Books were acquired in bulk, shipped to outside contractors, cut with commercial equipment, passed through high-speed scanners and then discarded or recycled.
The Judge Treated the Purchased and Pirated Books Differently
Judge Alsup divided Anthropic’s conduct into three legal questions. The first involved temporary and intermediate copies created while training language models. He found that the training process was transformative because the models were intended to generate new language rather than distribute substitute copies of the authors’ books.
The second question involved the PDFs created from lawfully purchased physical books. The court also found that specific process to be fair use. The judge emphasized that Anthropic had paid for each physical copy, destroyed the original while creating the digital replacement, retained the replacement internally and had not distributed the scanned PDFs outside the company. In the court’s view, Anthropic had created searchable and space saving replacements without increasing the total number of library copies.
The third question involved the permanent library of books downloaded from pirate sites. On that issue, Anthropic lost. The court rejected the argument that obtaining and retaining pirated copies became lawful merely because some might later be selected for transformative AI training. It also found that purchasing a legitimate physical copy afterward would not erase potential liability for an unauthorized digital copy acquired earlier.
The case was headed toward a trial over the pirate-library copies and possible damages before the parties reached a settlement.
First Sale Did Not Automatically Authorize the Scanning
The ruling is sometimes summarized as saying that once a company buys a book, it can do anything it wants with it. That is incomplete. The first-sale doctrine generally allows the lawful owner of a particular physical copy to resell, lend, alter or dispose of that copy. It does not ordinarily authorize making a new reproduction of the copyrighted work.
Anthropic’s right to destroy the paper books was therefore not the main legal controversy. The central question was whether it could reproduce their contents as digital files. Judge Alsup found that the scanning qualified as fair use under the particular circumstances before him. That was a case specific federal district court ruling, not a Supreme Court decision establishing that every owner may digitize any purchased book for any purpose.
The $1.5 Billion Settlement Covered Pirated Files, Not Destroyed Paper Books
On July 20, 2026, Judge Araceli Martínez-Olguín granted final approval to a $1.5 billion settlement in the Anthropic case. The settlement covered 482,460 works appearing in versions of LibGen and PiLiMi downloaded by Anthropic. The court estimated a recovery of approximately $3,000 per eligible work before fees and costs and described it as the largest copyright class-action settlement in history.
The agreement also required Anthropic to destroy original files downloaded from LibGen and PiLiMi, along with copies derived from those files, subject to evidence preservation requirements. That destruction provision applies to the pirated digital collections. It does not require Anthropic to delete the PDFs created from physical books it lawfully purchased and destructively scanned.
Meta Used Shadow Libraries and BitTorrent
The record in Kadrey v. Meta Platforms shows that Meta considered books particularly valuable for training its Llama models because books contain lengthy, edited and professionally structured writing. Meta initially explored licensing, and its head of generative AI discussed spending as much as $100 million. The company encountered difficulty determining who controlled AI training rights, which can remain with individual authors or vary by region.
In October 2022, Meta downloaded LibGen to evaluate whether its contents would improve model training. The court found that after licensing efforts failed and the matter was escalated to CEO Mark Zuckerberg, Meta decided in spring 2023 to use LibGen books as training data. Meta later downloaded Anna’s Archive, which included material from LibGen, Z-Library and other collections.
Meta used BitTorrent to obtain the collections. Torrenting can involve downloading pieces of files from multiple users while simultaneously uploading pieces to others. The parties agreed that Meta torrented LibGen and Anna’s Archive but disputed whether Meta uploaded copyrighted books to other users and, if so, how much. A Meta engineer wrote a script intended to prevent continued seeding after downloads were completed, although the parties disputed whether uploading occurred while downloads were underway.
The court found that Meta added books from the downloaded collections to datasets used to train Llama. All 13 named plaintiffs’ books appeared in the collections, and the court counted at least 666 copies associated with those plaintiffs across the relevant datasets. That figure represented copies tied to the named authors, not Meta’s total book collection.
Judge Vince Chhabria granted Meta summary judgment on those authors’ reproduction claims but limited the significance of the decision. The plaintiffs had failed to present sufficient evidence showing that Meta’s copying would harm the market for their books. The ruling did not establish that using copyrighted materials for AI training is generally lawful. The separate question of whether Meta infringed distribution rights through torrent uploads was not resolved by that fair use decision.
OpenAI Used LibGen-Derived Datasets for GPT-3 and GPT-3.5
A November 2025 federal discovery order in the consolidated In re OpenAI Copyright Infringement Litigation provides the clearest public account of OpenAI’s early book datasets. The order says it was undisputed that an OpenAI employee downloaded pirated books from LibGen in 2018. Plaintiffs contend those files became datasets initially called LibGen1 and LibGen2 and later renamed Books1 and Books2.
An OpenAI corporate representative testified about how the datasets were downloaded and said they were used to train GPT-3 and GPT-3.5. OpenAI stopped using the datasets for training in late 2021 and deleted them in mid-2022, approximately one year before the first lawsuits in the consolidated litigation were filed. According to OpenAI, Books1 and Books2 were the only training datasets it had ever deleted.
OpenAI initially said the deletion occurred because the datasets were no longer being used. It later withdrew that explanation and asserted that the reasons were protected by attorney-client privilege. The court examined OpenAI’s changing representations and partially granted the plaintiffs’ request for additional disclosure. The order also says copies of Books1 and Books2 were eventually recovered.
Microsoft and NVIDIA Disclosed Their Use of Books3
Microsoft and NVIDIA jointly developed the 530-billion-parameter Megatron-Turing Natural Language Generation model, known as MT-NLG. Their own technical publication identified Books3 as one of 15 datasets used in training. Books3 contributed 25.7 billion tokens, received a 14.3 percent sampling weight and was processed for approximately 1.5 training epochs.
That is direct company documentation, not an inference based solely on model outputs or allegations in a lawsuit. The Anthropic court order described Books3 as a collection of 196,640 books assembled from unauthorized copyrighted copies.
Authors later sued Microsoft in Bird v. Microsoft, alleging that its use of Books3 infringed their copyrights. The lawsuit was filed in federal court in New York in June 2025, but no final judgment has determined whether Microsoft’s use was infringing or protected as fair use.
NVIDIA is also defending separate litigation concerning Books3 and other alleged training sources. In May 2026, a federal judge allowed portions of Nazemian v. NVIDIA to proceed beyond the motion-to-dismiss stage. That means the allegations were sufficiently pleaded to continue; it does not mean NVIDIA has been found liable.
Google Scanned Millions of Books but Did Not Operate the Same Type of Program
Google presents the closest historical comparison to Anthropic because it operated one of the world’s largest physical-book scanning programs. Beginning in 2004, Google received books from participating research libraries, scanned the pages, performed optical character recognition and retained digital copies for Google Books. The system allowed users to search the texts and view limited snippets, while partner libraries could obtain digital copies of books from their own collections.
By the time of the 2015 appellate decision in Authors Guild v. Google, the company had scanned and indexed more than 20 million books. The Second Circuit Court of Appeals found the search and snippet system to be transformative fair use because it served research and discovery functions and did not provide the public with a meaningful substitute for purchasing the books.
The Google Books case documents mass scanning of physical books supplied by libraries. It does not describe Google buying millions of used books, removing their bindings and throwing away the originals.
A separate lawsuit filed against Google in July 2026 alleges that the company used copyrighted works to develop LaMDA, PaLM, Bard and Gemini. The complaint claims Google repurposed digital copies it possessed through Google Books, Google Play Books and other services and used web-crawl data containing material copied from piracy sites.
Those claims remain allegations. No court has yet determined that Google used the books as alleged, that the conduct was infringing or that Google’s earlier permission and fair use rights for book search extended, or did not extend, to AI training.
What the Evidence Establishes
As of late July 2026, Anthropic remains the only major AI developer for which the reviewed public court record establishes a large program based on buying physical books, cutting apart their bindings, scanning pages and discarding the originals.
For the other major companies, the strongest documented facts concern digital materials. Meta downloaded and torrented LibGen and Anna’s Archive, then added books from those collections to Llama training datasets. An OpenAI employee downloaded LibGen books, and OpenAI testimony says the resulting datasets were used for GPT-3 and GPT-3.5. Microsoft and NVIDIA publicly disclosed that their joint MT-NLG model was trained partly on Books3. Google scanned more than 20 million library books for Google Books, but that documented project was not an Anthropic-style purchase-and-destruction operation. A pending lawsuit now alleges that Google later repurposed digital copies for Gemini.
The verified story is already significant without extending it beyond the evidence: one leading AI company built a permanent private digital library by destroying millions of purchased books, while several other developers built or trained models using enormous digital book collections whose acquisition and legal status are now being examined in federal courts.
Is There an Obligation to Protect Physical Books?
Anthropic’s physical book operation relied partly on established used book retailers. Unsealed court records show the company purchased books in batches that sometimes reached tens of thousands of copies from sellers including Better World Books and the United Kingdom based World of Books. The records do not establish that every retailer knew the books would be cut apart and scanned.
A broader supply network now appears to be developing. ISBNdb, a book data and sourcing company, openly advertises bulk acquisition services for AI developers, with orders ranging from 1,000 to as many as one million titles. Its marketing emphasizes older, out of print and specialized books that may not already exist in a usable digital form, while offering confidentiality to its customers.
Independent booksellers in Europe have also reported unusual orders for seemingly unrelated and obscure titles. A Dutch dealer received a request linked to a company identifying itself as 2077AI for approximately 3,000 specific editions, while sellers in Germany, Switzerland and Spain described similar purchasing patterns.
