Have you noticed that older books are getting harder to find? A new investigation using a hidden tracking tag shows that Amazon is buying up huge shipments of rare books, scanning the pages, and then physically trashing them. The company uses this fresh, human-written text to feed its large computer models. By feeding the machines clean data, developers prevent the models from putting out bad information.
Trackers follow the books to a Las Vegas sorting center
Journalists worked with a rare bookseller to see who was buying all their inventory. They hid a tracking tag inside one rare book that was part of a large order. The book shipped across the country until it finally arrived at a specific company facility in Nevada.
Workers at this location receive heavy shipments of printed books. To move fast, they chop off the book bindings and feed the loose pages straight through high speed scanners. Once the machines turn the text into data, the leftover paper is thrown in the garbage. The company does not save these scans for humans to read, but rather to train artificial intelligence programs.
Don’t miss the best of The Mac Observer
Set us as a preferred source and our Apple reporting ranks higher in your Google Search results and Discover feed — one tap, no account changes.
Machines need human text to keep working the right way
Digital data is running dry. When an AI program learns from text made by another computer program, it eventually starts to break down. This causes the system to spit out boring, broken, and useless answers. To keep things running smoothly, developers have to feed the algorithms real text written by humans.
Antique books are packed with this kind of clean data. Other big tech companies are doing the very same thing to train their own systems. Judges say this is legal as long as the companies do not sell the scanned copies. Still, it means that many rare physical books are gone forever just to help a computer learn how to talk.
This is unacceptable. Someone who can need to stop this.