AI companies’ search for cleaner training data is spilling into the antiquarian and secondhand book trade.

Citing a recent Fortune report, TechFlowPost said rare-book dealers in the Netherlands, Germany, Switzerland and Spain have been receiving unusually large orders this year. Buyers have been asking for thousands of academic titles at a time, often across unrelated subjects, which dealers said did not resemble normal purchasing behavior.
One example involved de Vries, a bookseller in Haarlem, the Netherlands. He received an email from a Singapore company called 2077AI saying it was running a multilingual book acquisition project. Attached was an Excel file listing more than 3,000 books across different subjects, each with a publication identifier, along with a request for a quote and shipment to China. de Vries initially dismissed it as absurd and did not respond. Weeks later, after being contacted by a Dutch journalist, he realized the message was tied to a much larger appetite from the AI industry for training material.
Fortune said the company behind the purchases declined to answer questions about who would use the books and what models they were meant to train.
The article says 2077AI was not the only case. Another company, Canada-based Zoom Books, was described as placing bulk orders on European booksellers’ online stores between 3 a.m. and 5 a.m. every day, focusing on obscure academic books with little connection to one another. Some sellers told media outlets that shipping costs were higher than the books themselves, but buyers appeared unconcerned.
These purchases were not meant for reading. AI firms want printed books published before 2022 because those texts are seen as free of contamination from AI-generated content and therefore cleaner for large language model training. Once the books arrive, the standard process described in the report is destructive scanning: cut off the spine, scan every page at high speed, then destroy the physical copy.
Anthropic case exposed how book data was sourced
The clearest public record of this process came from a U.S. copyright lawsuit.
In August 2024, writers including Andrea Bartz sued Anthropic, the parent company of Claude, alleging that their works had been used without authorization to train AI models. By early 2026, more than 4,000 pages of Anthropic internal documents had been unsealed in court, and The Washington Post published a detailed account based on those records.

According to the documents, Anthropic downloaded more than 7 million books from the pirate ebook sites LibGen and Pirate Library Mirror via BitTorrent in 2021 and 2022, then stored them on internal servers. The records quoted management as viewing the process of negotiating licenses with publishers one by one as “too cumbersome.”
The three writers who filed the original complaint were only the first plaintiffs. The case expanded into a class action covering about 500,000 works. In June 2025, Judge Alsup issued a split ruling: using books to train AI counted as fair use, but downloading and permanently storing copies from pirate sites did not. That left the training act itself on one side of the legal line and the acquisition method on the other.
Anthropic faced a theoretical damages ceiling of $150,000 per book. Applied to roughly 500,000 works, that implied an enormous upper bound. The dispute was eventually settled for $1.5 billion, or about $3,000 per work on average. Final court approval came on July 20 this year, making it the largest class-action copyright settlement in U.S. history by amount, according to the article.
The “Panama Project” and destructive scanning
The story did not end with pirated digital libraries.
TechFlowPost cited a January 2026 report by The Washington Post saying that the unsealed Anthropic documents also described another initiative. In early 2024, Anthropic launched an internal effort code-named the “Panama Project.” The company brought in a former head from the Google Books era and began buying printed books in bulk from secondhand-book platforms, with a plan to process between 500,000 and 2 million books within six months.
The workflow was straightforward: buy the books, cut off the spines, scan them and destroy the originals. Internal records described the effort as “destructive scanning of all the books in the world” and added, “we do not want the world to know we are doing this.”
The difference from downloading pirate copies was that the books in this project were purchased. A judge later found that process lawful because one purchased copy was converted into one digital copy and the physical original was destroyed, so there was no extra act of reproduction under copyright law.
In practical terms, the report argues, that decision marked out a route the rest of the industry could follow. If pulling books from pirate libraries invites copyright claims, then buying legal copies, scanning them and disposing of the originals lowers the legal risk.

Supply infrastructure is starting to form
That route is no longer limited to one company.
According to Decrypt, federal judges in separate copyright cases involving OpenAI and Meta also issued similar fair-use rulings. TechFlowPost said the result is that “book burning for AI training” is shifting from an isolated case to a practice receiving broader support through case law.
Where there is demand, suppliers follow. Citing 404 Media, the article said ISBNdb, a book database company with data on more than 110 million books, had posted a marketing page offering “paper book corpus procurement services” for AI companies. The page advertised bulk sourcing from 1,000 books to 1 million books per order and said strict confidentiality agreements would protect buyers’ identities.
After the report was published, ISBNdb removed the page and posted a statement on its website saying the service “was never actually launched” and that the page was only meant to “test market interest.”
Still, 404 Media cited an anonymous bookseller who said weekly sales had jumped from about 20 books to several hundred starting in April this year. Buyers were placing precise orders by ISBN and choosing obscure titles with no relation to one another. The one shared feature was that every book had been published before 2022.
That date matters because, as the article notes, the wording on ISBNdb’s now-removed page pointed to the core selling point: printed books published before 2022 contain no AI-generated content and can be marketed as cleaner training material.
The supply chain sketched in the report looks like this:
- database providers supply title lists and anonymous procurement channels;
- secondhand booksellers around the world supply the physical books;
- scanning firms cut off spines and digitize pages at high speed;
- disposal companies handle the remains of the physical books;
- the resulting digital text is fed into AI training pipelines.
Criticism and a slower alternative
The process has also drawn criticism.

On July 27, Elon Musk posted on X criticizing what the article described as a “modern book burning” style of industrial workflow. He wrote: “I have asked the SpaceXAI team to preserve any rare books they process in a library and to scan them the ‘dumb way’ (non-destructively), rather than simply cutting off the spine and scanning.”
Without court disclosures, the names of buyers might never have surfaced at any point in the chain, and the public might not have seen how procurement, scanning and disposal were tied together.
The article also notes that destruction is not the only technical option. Google scanned more than 40 million books for Google Books using patented non-destructive scanning technology, then returned the books to libraries rather than destroying them. In July this year, Musk also said publicly that xAI should use non-destructive scanning for rare books and preserve the originals in a library.
That alternative is feasible, just slower and more expensive. TechFlowPost cited a now-deleted ISBNdb blog post that argued destructive scanning became the industry’s preferred approach for one reason: lower cost and higher speed. Non-destructive scanning requires specialized equipment and more labor, and processing time per book is several times longer.
In a competitive race for AI capability, companies that secure more high-quality text earlier may gain an edge. Court rulings, as described in the article, have not blocked the buy-scan-destroy model.
But the report closes on a different point. A printed book can be resold, borrowed and passed down. It can sit for 20 years in a secondhand shop and still be discovered by a stranger. Once scanned and destroyed, it becomes private training data sitting on a company server, no longer part of a public circulation of knowledge.
The author wrote that he uses AI every day to draft articles, check information and organize ideas, and acknowledged that these tools are changing how many people work. Even so, the piece argues that looking closely at how some training data is acquired leaves an uneasy picture: each step may have a business logic and a legal rationale on its own, yet taken together, it raises harder questions.

