The short version
An investigation tracked a rare book to an Amazon facility where it was reportedly destroyed to scan its text for training AI models.
An investigation from 404 Media followed a rare book with a tracking device to an Amazon facility in Las Vegas. The report claims Amazon buys rare texts, removes their spines, and scans them to collect training data for its large language models. These pre-2022 books are prized because they contain human-written text not available online, which helps stop AI models from degrading.
Key takeaways
- An investigation followed a rare book to an Amazon facility called VGT3 in Las Vegas.
- Amazon reportedly buys rare books and damages them by cutting off spines for efficient scanning.
- The scanned text becomes training data for Amazon’s large language models (LLMs).
- Pre-2022 books are highly desired because they provide text not found online and are certain to be human-written.
- Using human-written text helps avoid “model collapse,” a quality drop that happens when LLMs learn from AI-generated content.
The Investigation: Tracking a Rare Book to an Amazon Facility
404 Media placed a tracking device inside a rare book to see where it went. The report shows the book was delivered to an Amazon facility in Las Vegas.
Workers know this location as VGT3. It uses a unique symbol: a dinosaur holding a book in its claws.
Purpose of the Acquisition
Amazon says it buys books through standard sales channels to better its products and services. The investigation suggests the company acquires rare texts, slices off their spines, and scans them. This content is needed for training large language models, which demand huge amounts of text data.
Rare books are especially useful for AI training because they contain text you cannot find on the internet. Also, books printed before 2022 are definitely not written by an AI. This helps prevent “model collapse”—a decline in output quality that can happen when LLMs train on AI-generated content.
The Process: Destroying Books for Data Scraping
Amazon buys many rare books through commercial sellers. A 404 Media report states the company gets these texts to improve its products and services.
Destructive Scanning for AI Training
This method requires cutting the spines off these rare books for quick scanning. Investigators discovered this after tracing a tracked book to the VGT3 Amazon facility in Las Vegas. The scanned text from these books then trains Amazon’s large language models.
These rare and out-of-print books are a highly desired training source because they offer text not available on the internet. This information is very useful for AI training. Texts published before 2022 ensure they were not written by an LLM, helping to stop model collapse that can result from training on AI-made content.
The Motivation: The Scarcity of High-Quality, Pre-2022 Text
Firms like Amazon need enormous amounts of text data to train their large language models. A report notes these models have already consumed most available internet text, creating demand for new training sources.
Rare, out-of-print books have become a sought-after new source simply because they are not online. These texts provide a new supply of data that AI has not yet used for training.
Avoiding Model Collapse
Books printed before 2022 have particular worth. They could not have been written by an LLM. This matters because when LLMs train on AI-generated text, they face “model collapse,” a situation where the AI’s output quality falls. Using pre-2022 books guarantees the data is human-made, removing this contamination risk.
Corporate Justification and Industry Context
Amazon gave a clear reason for its actions in a statement to 404 Media, saying it “purchases books through commercial channels to improve the products and services customers use.” This practice of buying physical books is described as a standard purchase effort to collect data for improving its AI models.
This action fits a wider industry setting where AI firms actively hunt for large text volumes. As the report mentions, large language models have already consumed much of the text available online. This has pushed some companies to look for other sources. The source specifically notes that, for Anthropic, this meant using illegally pirated books for training data.
Rare and out-of-print books are particularly sought after because they represent a fresh, unused collection of text. Also, texts published before 2022 are considered useful for AI training because they pose no risk of being AI-generated. This helps avoid “model collapse,” which can lower an AI’s output quality.
📡 Original reporting: Hacker News · AI. AI Craft Technologies’ news engine summarised and rewrote this story in our own words; facts are drawn from the linked source.
⚙️ How this article was made — fully automated
This is a live demo of the ACT News Factory engine. Want one running on your own site? See our services →



