Robots Atlas>ROBOTS ATLAS
Artificial Intelligence

Amazon is destroying rare books to train AI models

Sir Robot21 August 2026 · 3 min read
Amazon is destroying rare books to train AI models

404 Media has found that Amazon buys rare, often offline books, cuts off their spines and scans the pages to obtain training data for AI models. The reporters tracked one shipment by hiding an AirTag inside it — it ended at an Amazon facility in Las Vegas. It shows how far AI companies will go for data after exhausting the public internet.

Key takeaways

  • 404 Media tracked a book shipment with an AirTag to an Amazon facility.
  • Destination: the VGT3 section of the LAS8 facility in northeast Las Vegas.
  • VGT3 "destructively scans large volumes of books" — confirmed by workers.
  • An order of about 1,000 books placed through Biblio by an anonymous buyer.
  • Amazon's official statement: it buys books "through commercial channels to improve products and services".

How the practice was uncovered

It started with an unusual purchase. A bookseller received an order for roughly 1,000 books from an anonymous, price-insensitive buyer — the typical profile of an AI company.

404 Media placed an Apple AirTag in one volume. The signal led to the VGT3 section of Amazon's LAS8 facility in Las Vegas. According to the outlet, VGT3 "destructively scans large volumes of books," a practice workers confirmed in forum discussions. Above the entrance sits a logo of a dinosaur holding a book in its claws.

Why rare books specifically

It is about data that is not online. Language models have exhausted easily available internet text, so companies turn to rare, out-of-print titles.

The date matters: texts predating 2022 were written before the era of mass AI-generated content, which helps avoid "model collapse: The quality decay of an AI model trained on data generated by other models — errors compound with each generation." — the degradation of a model fed its own synthetic data. Physical books thus become a source of "clean" training data.

Not the first book dispute

Sourcing books to train AI already has a legal history. The article recalls the case against Anthropic, accused of using pirated books — settled for $1.5B in July 2026. The difference matters: Amazon says it buys through legal channels. The controversy remains the irreversible destruction of rare, sometimes unique copies.

$1.5BAnthropic's settlement over training on pirated books (July 2026).TechCrunch

Why it matters

The case shows the real cost of AI's data hunger. When the public internet no longer suffices, physical rare texts become valuable — and acquiring them means irreversibly destroying copies.

That creates tension between AI companies' interests and the protection of cultural heritage and the antiquarian market. Legality of purchase does not settle the ethical question: whether training models justifies the permanent loss of rare books no one will reprint.

What's next?

  • Amazon has not disclosed the scale of the operation or the number of scanned books — pressure for transparency will grow.
  • Booksellers and libraries may tighten rules on large, anonymous orders.
  • The case will feed the regulatory debate on training-data provenance, after the Anthropic settlement.

Sources

Share this article