GKRootWire
AI Google Adds 'Preferred Source' Button to Help Publishers Fight AI Traffic LossesGadgets Linkdaze Launches a Smart Calendar Aimed at Running Your Whole HouseholdSecurity Popular Rust Crate arrayref Hijacked to Spread Infostealer MalwareCloud & Sysadmin GitHub Details Cause of August 17 Outage, Outlines Reliability FixesDev Tools Show HN: 'Huzzah' Proposes a Fresh Take on AI-Assisted CodingCloud & Sysadmin The Weird Science of Cooling Data Centers With UrineAI Google Adds 'Preferred Source' Button to Help Publishers Fight AI Traffic LossesGadgets Linkdaze Launches a Smart Calendar Aimed at Running Your Whole HouseholdSecurity Popular Rust Crate arrayref Hijacked to Spread Infostealer MalwareCloud & Sysadmin GitHub Details Cause of August 17 Outage, Outlines Reliability FixesDev Tools Show HN: 'Huzzah' Proposes a Fresh Take on AI-Assisted CodingCloud & Sysadmin The Weird Science of Cooling Data Centers With Urine
AI

Amazon Is Reportedly Scanning and Destroying Rare Books to Feed AI Models

With the open internet largely tapped out, Amazon is turning to physical rare books as fresh training data for its large language models.

Amazon has reportedly been acquiring rare and out-of-print books, scanning their contents, and destroying the physical copies afterward, all to gather fresh text for training its AI models. The logic is straightforward: most publicly available online text has already been scraped and used, so unique, never-digitized books offer a rare source of novel language data.

Critics argue this process is troubling because it permanently destroys physical artifacts, some of which may be irreplaceable, in service of feeding a data-hungry AI pipeline. It also raises questions about transparency, since there's no public catalog of what's being scanned or destroyed, and no guarantee the content will ever benefit anyone beyond Amazon's own models.

The practice highlights just how strained the supply of fresh training text has become for large AI labs.

Why it matters: As internet-scraped data dries up, AI companies are increasingly hunting for untapped text sources, and physical archives are apparently fair game even at the cost of destroying originals. This raises real preservation and accountability concerns, especially if there's no public record of what's lost in the process.

Sources: TechCrunch