Amazon Turns to Physical Rare Books for AI Training Data
The e-commerce giant is acquiring out-of-print volumes and scanning them to feed its language models, raising questions about how far tech companies will go to secure training material.

The Digital Pivot to Physical Archives
Amazon has begun acquiring rare and out-of-print books, removing their spines, and digitizing them to train artificial intelligence models. The operation centers on a Las Vegas facility that uses a dinosaur-clutching-book symbol as its identifier, a curious choice for a company whose founding narrative centered on democratizing access to literature.
The practice surfaced after an investigation involving a tracking device placed inside a rare book, which eventually arrived at Amazon's VGT3 facility in Nevada. Amazon confirmed the program, stating it "purchases books through commercial channels to improve the products and services customers use," though the company declined to elaborate on the scale or specific criteria for acquisition.
Why Pre-Digital Text Matters
The move reflects a fundamental challenge facing large language model developers: they're running out of training data. After scraping publicly available internet content and, in some documented cases, incorporating illegally obtained digital book collections, AI companies now face diminishing returns from online sources.
Physical books, particularly those published before the widespread adoption of generative AI in 2022, offer two distinct advantages. First, they represent text that never existed in easily accessible digital form, expanding the corpus beyond what web crawlers can reach. Second, and perhaps more critically, pre-2022 publications guarantee human authorship at a time when LLM developers worry about model collapse.
Model collapse occurs when AI systems train on outputs generated by other AI systems, creating a feedback loop that degrades performance over time. Think of it as a photocopy of a photocopy: each generation loses fidelity. For companies betting billions on AI capabilities, sourcing verifiably human-written text has become a competitive necessity.
The Irony of Destruction
There's a particular tension in watching Amazon, which built its empire by making books more accessible, now destroying physical copies to feed algorithms. The company's original vision promised that anyone with internet access could obtain nearly any book ever printed. Two decades later, that same infrastructure supports the systematic dismantling of rare volumes that may exist in limited quantities.
At DailyTechWire, we've tracked the AI industry's data hunger across the region, from Seoul's KT Corporation building Korean-language corpora to Singapore's AI startups licensing Southeast Asian literary archives. But the physical destruction of rare books marks a new threshold, one where the raw material isn't bits but bound paper with potential cultural and historical value.
The practice raises questions beyond mere symbolism. What happens when a unique or near-unique edition gets destroyed for training data? Who decides which books are expendable? Amazon's statement about "commercial channels" suggests these are legally purchased items, but legal acquisition doesn't necessarily align with preservation ethics.
The Broader Data Scarcity Problem
Amazon's book-scanning operation sits within a larger industry reckoning. Major AI labs have already exhausted much of the accessible internet. OpenAI, Google DeepMind, and Anthropic have ingested Common Crawl datasets, Wikipedia in dozens of languages, GitHub repositories, academic papers, and Reddit discussions. The low-hanging fruit is gone.
This scarcity has sparked increasingly aggressive data strategies. Some companies now license content directly from publishers and media organizations, paying for access to archives. Others have faced legal challenges: Anthropic currently defends itself against claims it trained models on pirated book collections. Still others, like Amazon, are apparently going analog.
The shift to physical books as training material also highlights geographic inequities in AI development. Rare book markets concentrate in major Western cities and established Asian hubs like Tokyo and Hong Kong. Texts from regions with less robust preservation infrastructure or smaller antiquarian markets may simply vanish from the training data pool, potentially skewing model outputs toward already over-represented perspectives.
What Amazon Gains
For Amazon, the rare book program serves multiple strategic purposes. The company operates its own AI division, Amazon Web Services offers foundation model hosting, and Alexa relies on natural language processing. Access to unique training data could differentiate Amazon's models in an increasingly crowded field.
The Las Vegas facility's existence also suggests this isn't a small-scale experiment. Purpose-built infrastructure for book processing indicates volume and continuity. Amazon's logistics network, built to move millions of physical books daily, now enables the reverse operation: funneling rare volumes into a centralized scanning operation.
The dinosaur symbol at VGT3 might be more apt than intended. Just as dinosaurs went extinct while giving rise to new forms, physical rare books are being sacrificed to birth the AI systems that may eventually replace much human writing. Whether that's evolution or extinction depends largely on where you sit.
Unanswered Questions
Amazon hasn't disclosed how it selects books for acquisition, what happens to the physical remains after scanning, or whether it maintains any obligation to preserve digital copies for non-commercial purposes. The company also hasn't addressed whether it's targeting specific subject areas, languages, or publication eras.
There's no indication Amazon is breaking laws. Purchasing books and scanning them for internal use likely falls within existing copyright frameworks, especially for out-of-print works. But legality and wisdom don't always align. The AI industry's data demands are creating incentives that previous generations of technologists never faced, and the regulatory environment hasn't caught up.
As the foundation model race intensifies, expect more companies to pursue unconventional data sources. Amazon's rare book program won't be the last example of AI development colliding with preservation values. The question isn't whether tech giants will keep searching for training data; it's what else they'll be willing to dismantle to find it.


