The Race to Scan Books for AI Training Is Destroying Physical Copies
Independent booksellers report unusual bulk purchases of rare and out-of-print volumes, followed by their likely destruction in the rush to feed language models.

A Quiet Pattern Emerges in the Book Trade
Independent booksellers have begun noticing a troubling pattern: bulk orders for older books, sometimes rare or out-of-print editions, purchased with little regard for condition or collectibility. The buyers show no interest in preservation, resale value, or even whether the spines remain intact. What matters is the text inside, and how quickly it can be extracted.
The practice points to an uncomfortable reality in the AI training pipeline. Language models need vast quantities of long-form, high-quality writing to improve their output. Academic papers and web scrapes only go so far. Books, particularly older works that sit outside active copyright enforcement or have fallen into obscurity, represent a rich seam of narrative complexity, varied syntax, and domain knowledge that newer datasets cannot replicate.
At DailyTechWire, we've tracked the evolution of training data strategies across major AI labs, and the shift toward physical books has been building for over a year. What was once a niche operation involving library partnerships and careful digitization has morphed into something far more industrial.
The Economics of Destruction
Scanning a book properly, the kind of work that archives and libraries undertake, is slow. It requires careful handling, page-by-page photography with specialized equipment, and post-processing to correct for curvature and lighting. A single volume can take hours.
The alternative is faster and cheaper: remove the spine with a blade or industrial cutter, feed individual pages through a high-speed scanner, then discard the remains. For companies operating under venture timelines and competitive pressure, the choice is obvious. Speed trumps sentimentality.
The financial incentives align in troubling ways. A rare book that might sell for fifty or a hundred dollars to a collector holds a different kind of value when its contents can be tokenized and fed into a model that powers a billion-dollar valuation. The physical object becomes disposable once its informational payload has been extracted.
Booksellers report that many of these bulk buyers operate through intermediaries or use generic company names that obscure their ultimate purpose. Payment is prompt, questions are few, and there is no interest in return customers or relationship-building. The transaction is purely extractive.
What Gets Lost
The concern is not just about individual volumes, though the loss of any book is worth lamenting. The deeper issue is that some editions, particularly those printed in small runs or from regional publishers, may exist in only a handful of physical copies worldwide. Once destroyed for scanning, they are effectively erased from the material record.
Digital copies are not perfect substitutes. They lack marginalia, inscriptions, and the physical traces of how a book moved through the world. They cannot convey the texture of paper, the smell of aging glue, or the weight of an object that survived decades or centuries. These qualities matter to scholars, collectors, and anyone who believes that books are more than vessels for text.
There is also a question of access. If a rare book is scanned and then destroyed, and the resulting digital file is locked inside a proprietary AI training corpus, the public loses twice. The physical object is gone, and the digitized version remains inaccessible except as an abstracted influence on model weights.
Alternatives Exist, But They Are Slower
The frustration among preservationists is that this destruction is avoidable. Non-destructive scanning technology exists and is used regularly by libraries and archives. Overhead scanners with cradles that support book spines can digitize volumes without removing a single page. The process is slower, but it leaves the book intact for future readers.
Several large-scale digitization projects, including those undertaken by research institutions and national libraries, have proven that it is possible to build massive text corpora without sacrificing the source material. The difference is that these projects operate on timelines measured in years, not quarters, and prioritize preservation alongside access.
For AI companies operating in a competitive environment, that patience is a luxury. The race to build the next generation of models is measured in months, and training data is the fuel. If acquiring that data means buying books in bulk and feeding them into industrial scanners, the calculation is straightforward.
The Regulatory Vacuum
There is currently no legal framework that prevents this practice. Buying a book confers ownership, and what the buyer does with that book afterward, short of violating copyright in the use of its contents, is largely unregulated. If the text is in the public domain, or if the buyer believes they can argue fair use or equivalent doctrines in their jurisdiction, destruction of the physical copy carries no penalty.
Some jurisdictions have heritage laws that protect culturally significant objects, but these rarely extend to books unless they are exceptionally rare or historically important. A first edition of a novel by a mid-tier author from the 1960s, or a technical manual from the 1980s, would not typically qualify for protection even if only a few copies remain in circulation.
The result is a regulatory vacuum in which market forces alone determine the fate of physical books. If an AI company is willing to pay more than a collector, the book goes to the scanner.
A Cultural Reckoning
The broader question is whether society is comfortable with this trade-off. The benefits of advanced AI models are real and widely distributed. Better language understanding can improve translation, accessibility tools, education, and countless other applications. But if the cost of that progress is the permanent loss of physical books, including rare and irreplaceable editions, the calculation becomes less clear.
There is a romantic attachment to books that transcends their utility. They are objects of beauty, markers of intellectual history, and tangible links to the past. Watching them reduced to raw data feels like a category error, a failure to recognize that not everything can or should be optimized for machine learning.
At DailyTechWire, we've observed similar tensions across other domains where AI development collides with cultural preservation. The pattern is consistent: the technology moves faster than the institutions meant to steward the material it consumes. By the time preservation advocates raise alarms, the practice is already entrenched and scaled.
What Booksellers Are Doing
Some booksellers have begun refusing bulk orders from buyers they suspect are acting on behalf of AI companies. Others are implementing informal vetting processes, asking questions about intended use and declining sales when the answers are evasive. It is a small-scale resistance, and it cannot stop the broader trend, but it reflects a desire to retain some agency over where these books end up.
There is also a growing conversation within the rare book trade about collective action. Industry groups are discussing whether to establish guidelines or voluntary standards that would discourage sales to known destructive buyers. The challenge is enforcement: the market is decentralized, and individual sellers operate under their own economic pressures.
Meanwhile, some institutions are accelerating their own digitization efforts in the hope of creating publicly accessible copies before physical books disappear into private training corpora. These projects are under-resourced and cannot match the pace of commercial scanning operations, but they represent an attempt to preserve both the text and the object.
The Path Forward
If the destruction of books for AI training is to be curtailed, it will require a combination of regulatory intervention, industry self-regulation, and shifts in corporate behavior. Governments could extend heritage protections to cover a broader range of books, or require that any entity engaging in bulk scanning make the resulting digital copies publicly available. Industry groups could establish norms that discourage destructive practices and reward preservation-minded digitization.
AI companies themselves could choose to prioritize non-destructive methods, even if it slows their timelines. Some already do, partnering with libraries and using overhead scanners to build training datasets without destroying source material. These efforts deserve recognition, and they demonstrate that the race for training data does not have to leave a trail of shredded spines in its wake.
For now, the practice continues, driven by the same forces that shape much of the AI industry: competition, capital, and the belief that speed is survival. Whether that belief is correct, and whether the cost is worth paying, remains an open question. But for anyone who has ever held an old book and felt the weight of the centuries in their hands, the answer is already clear.


