As the demand for high-quality text for AI model training increases, some companies have been exposed for purchasing large quantities of paper books through intermediaries, disassembling and scanning them to import training data, and then discarding the original books. This practice is driving up the demand for secondhand books and raising concerns about the permanent depletion of rare books.
Anonymous book purchase by middlemen
According to foreign media reports, some suppliers have begun offering bulk book-finding services to AI clients, capable of collecting hundreds of thousands of books in a short period of time, and promising confidentiality for buyers. Due to ongoing scrutiny of the sources of training data, such purchases are typically not made directly in the name of the AI company.
Market demand is primarily focused on works published before the widespread adoption of generative AI. This type of content is generally considered a more stable sample of human writing, suitable for model training. The report also mentions rising demand for niche books, out-of-print books, and foreign language books.
The demand for secondhand books is rising.
One bookseller said that after AI buyers entered the market, their weekly sales rose from about 20 books to several hundred. The sales increase brought direct benefits and also helped them clear previously unsold inventory.
However, he also stated that he does not agree with the final use of these books, and is particularly worried that some uncommon or out-of-print books will be destroyed directly after scanning, leading to a further reduction in physical copies.
- Books in increasing demand include niche books, out-of-print books, and foreign language books.
- Some sellers reported a significant increase in sales in a short period of time.
- The crux of the controversy lies in whether the original book is preserved after scanning.
Court rulings and copyright lawsuits proceed in parallel.
This approach is similar to Anthropic's previous large-scale book digitization project. Last summer, a federal judge in San Francisco ruled in Bartz v. Anthropic that if a book was legally purchased, converting it into a digital copy could still constitute "transformative fair use," even if the original is destroyed after scanning.
Subsequently, similar fair use rulings emerged in other copyright cases involving OpenAI and Meta. However, the legal risks surrounding the source of training data have not disappeared.
This week, another federal judge in the same district approved a $1.5 billion copyright settlement. Under the settlement, Anthropic must pay damages to thousands of authors, approximately $3,000 per book, because the company used pirated books to train Claude.
The issue of rare book preservation is gaining traction.
As the controversy grows, whether scanned physical books should be preserved is becoming a new point of discussion. The report mentions that Musk has publicly opposed the practice of directly disassembling and scanning books.
He stated on the X platform that he has asked the SpaceX AI team to preserve rare books and to scan them in a way that does not damage the spine, rather than disassembling the books and processing them quickly.











