Amazon Tracked Destroying Rare Books for Model Training
An AirTag inside a bulk book order led to a Las Vegas warehouse where teams dismantle volumes to scan pages, raising questions about training data sourcing.

The Tracker That Told the Story
For months, independent booksellers across North America noticed an unusual pattern: bulk orders for out-of-print titles, rare editions, and specialized volumes purchased at scale, then vanishing into logistics networks without clear destinations. The transactions felt different from typical library acquisitions or institutional purchases, and sellers began comparing notes. The hypothesis emerged quietly in online forums and at book fairs: someone was buying these volumes not to preserve or resell them, but to digitize and destroy them.
Proving it required something simple. A bookseller working with investigative journalists agreed to slip an AirTag into the spine of a rare book included in one such bulk order. The device pinged its location across the country until it settled at a warehouse in Las Vegas designated VGT3. The facility belongs to Amazon, and inside, according to reporting that tracked the device's journey, teams focus on a single task: separating books from their bindings and feeding pages through high-speed scanners.
The warehouse door displayed a logo showing a Tyrannosaurus rex mid-bite into a book, an image that now reads less as internal branding humor and more as a candid mission statement.
What the Facility Reveals About Training Pipelines
Amazon has not commented specifically on whether the scanned material feeds its AI models, offering only general statements about its operations. But the Las Vegas facility's existence aligns with what those of us tracking foundation model development have observed over the past eighteen months: a sharp escalation in the acquisition of non-digital text sources as publicly available web corpora begin to plateau in usefulness.
Training a frontier model today requires hundreds of billions of tokens. Web scraping, once the default method, now yields diminishing returns as the same content circulates across platforms and as publishers block crawlers. Books, particularly older or specialized texts not available in digital formats, represent a corpus that is linguistically rich, editorially curated, and in many cases, out of copyright or difficult to trace.
Rare books add another dimension. Editions with annotations, regional printings, technical manuals, and volumes published in limited runs contain language patterns and domain knowledge absent from mainstream digitized libraries. For a model aiming to handle nuanced queries across disciplines, these texts are high-value inputs.
The method Amazon appears to have chosen, destructive scanning, is faster and cheaper than preservation-grade digitization. Tearing a book from its spine allows pages to pass through sheet-fed scanners at speed, eliminating the need for careful handling or specialized equipment. The trade-off is irreversible: the book is gone.
The Economics of Bulk Acquisition
Booksellers report that bulk orders often come through intermediaries or procurement platforms, obscuring the end buyer. Prices offered are competitive but not extravagant, typically aligned with wholesale rates for used inventory. For sellers managing large estates or clearing warehouse stock, these orders are attractive, they move volume quickly without the friction of individual sales.
What makes the practice contentious is not the purchase itself but the intent. Books sold to libraries, archives, or collectors enter ecosystems where they remain accessible. Books sold for destructive scanning do not. Once digitized, the physical object holds no value to the buyer and is discarded.
This creates an asymmetry. The seller believes they are contributing to distribution or preservation. The buyer is extracting data from an artifact and eliminating the artifact. The transaction is legal, but the outcome, particularly when rare or hard-to-replace editions are involved, troubles the communities that steward these materials.
Why Rare Editions Matter Beyond Data
The argument Amazon and similar firms might make is straightforward: if a book is out of print, not digitized, and available for sale, purchasing it to create a digital copy serves a democratizing function. The text becomes machine-readable, searchable, and usable in applications that benefit millions.
But rare books are not just text containers. They are historical objects. Marginalia, binding choices, print quality, publisher notes, and distribution history all carry information that a raw text scan strips away. A 1920s technical manual on radio engineering, for example, tells you not only what engineers knew but how that knowledge was taught, illustrated, and disseminated in a specific moment.
Destroying such volumes to train a model that will generate approximations of technical language is a trade-off with long tails. Once the book is gone, future researchers studying the history of technology, pedagogy, or publishing lose access to the primary source. The model gains tokens. The archive loses an artifact.
Collectors and institutional buyers have begun adjusting. Some are pulling rare inventory from public listings. Others are requiring declarations of intent before completing sales. But the mechanisms to enforce this remain informal, and the scale of bulk buying outpaces the capacity of individual sellers to vet every transaction.
The Silence Around Sourcing
Amazon's refusal to address the AI training question directly mirrors a broader industry pattern. Firms building foundation models rarely disclose the composition of their training corpora in detail, citing competitive advantage and legal complexity. When pressed, they point to licenses, fair use arguments, or partnerships with content providers.
Rare books occupy a gray zone. Many are out of copyright, making their use technically permissible. Others are in copyright but difficult to trace to rights holders. Still others are sold secondhand under first-sale doctrine, which allows resale but does not explicitly authorize reproduction for commercial purposes.
The legal ambiguity creates operational space, but it also generates reputational risk. As the AirTag story circulates, Amazon now faces questions not about legality but about values: whether advancing model capability justifies erasing physical culture, and whether firms with the resources to digitize non-destructively should be held to that standard.
What Happens Next
The tracking revelation will likely accelerate three trends already underway. First, rare book dealers will tighten sales practices, adding friction to bulk transactions. Second, advocacy groups focused on digital rights and cultural preservation will push for disclosure requirements around training data sourcing. Third, competitors will scrutinize their own supply chains, aware that similar practices, if exposed, carry similar risks.
For Amazon, the immediate challenge is narrative control. The T-rex logo, intended as internal iconography, now functions as a symbol of algorithmic appetite, literalizing concerns that AI development consumes more than it creates. Whether the company addresses this directly or continues deflecting will signal how seriously it weighs public trust against training efficiency.
At DailyTechWire, we have tracked the collision between AI scale and archival ethics across multiple domains: medical imaging datasets scraped without consent, voice corpora built from call center recordings, code repositories ingested wholesale. The rare book case is distinct in one respect, it involves objects with singular physical presence. Once scanned and shredded, they do not exist elsewhere.
The question is not whether firms will continue sourcing hard-to-digitize materials. They will. The question is whether they will do so transparently, preserve what they scan, and recognize that training data is not ethically neutral simply because it is legally available. The AirTag in the spine of a rare book has made that question harder to ignore.


