After tainting natural resources in the name of AI, tech giants have now set their sight on books to train their new AI models.
Companies like Anthropic are buying and destroying millions of second-hand antique books through mediators, slicing their spines using industry grade hydrolic press, to then feed into high-speed scanners to train AI models.
To avoid training their models on AI-generated content—often referred to as “AI slop”—that has proliferated across the internet, artificial intelligence developers are quietly turning to physical books as a reliable source of high-quality training material. Companies are desperate to preserve human-authored knowledge in their new models, to be able to train the algorithm on a clean slate.
First reported by 404 media, the following revelations were exposed through unsealed court records and reports.
To acquire these massive print collections without drawing public backlash, AI firms are quietly employing commercial intermediaries.
READ: Meta tracks employee activity to train AI systems (April 21, 2026)
Book database operators and online marketplace sellers reported an unprecedented surge in bulk orders ranging from several thousand copies to a million volumes in a single purchase.
Independent booksellers who usually moved only a handful of books each week have admitted their order volumes have jumped exponentially overnight regardless of price, topic, or author.
While the huge demand for books by tech companies have profited the used book marketplace, the sellers aren’t entirely happy about how and why these books are being used.
“I’ve been well suited for these sales with inventory from overseas and foreign language books. On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped,” a seller told 404 media.
The critics allege that this practice is a part of Anthropic’s “Project Panama,” a digitization initiative hired scanning firms like Datamation Information Services to process massive batches of print volumes according to tom’s Hardware.
Once these shipments arrive at industrial scanning facilities, processing speed takes priority over historic preservation.
Rather than using slower overhead cameras to photograph pages without damage, contractors frequently opt for high-speed destructive scanning.
The process converts physical volumes into searchable digital text in seconds, leaving behind piles of shredded, unusable paper.
This practice has drawn intense legal scrutiny. Anthropic reportedly spent millions purchasing books from Better World Books for its Claude models, and while courts recognized AI training itself as fair use, the company was hit with a $1.5 billion penalty for storing seven million pirated titles.
Meanwhile, major book publishers have launched similar copyright lawsuits against Google over its use of copyrighted works to train Gemini models.
As the matter gains more attention and backlash, Elon Musk as an AI developer himself, took to X to speak against the practice stating that he has personally asked the “SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning.”
The internet observers, historians and book enthusiasts have expressed concern on social media, that once all the physical books are digitalized behind a corporate wall, the public will lose access to them. They also worry that this practice could eventually lead to the manipulation—or even erasure—of original historical records and cultural heritage, creating opportunities for falsehoods and conspiracy theories to be presented as authentic history to future generations.


