Amazon is buying rare books, removing their spines, and digitizing the pages to expand the material available for training artificial intelligence systems.
The operation includes out-of-print works or those difficult to find in digital format. These copies offer access to texts that are not available on the internet and can expand the dataset used in developing language models.
The practice came to light after a tracking device placed inside a rare book followed the copy to an Amazon facility in Las Vegas, known as VGT3. The location uses a dinosaur holding a book as its visual identifier.
Asked about the operation, Amazon said it “buys books through commercial channels to improve the products and services customers use.”
Why rare books are of interest to AI
Large language models depend on large volumes of text for training. Printed works that have never been digitized offer an additional source of content beyond the materials publicly available on the internet.
Old books also have another relevant characteristic for this process: texts published before the popularization of generative AI were produced by humans, without content created by large language models.
This type of material can be used to reduce the presence of synthetic texts in training sets. When models begin to repeatedly learn from content produced by other AIs, researchers point to the risk of a phenomenon known as “model collapse.”
In this scenario, the quality of results can deteriorate as successive generations of artificial content feed new training cycles.



