Intent not malicious, it's indifference: Thousands, if not millions, of books sent to AI woodchipper

François Picard is pleased to welcome investigative journalist Emanuel Maiberg, Co-founder of 404 Media. As we embark on AI's uncharted waters, Maiberg has made a most unexpected discovery:: the mass acquisition, dismemberment and scanning of rare books for use as training data. His investigation reveals a globalised and deliberately opaque supply chain in which books purchased through anonymous marketplaces can ultimately arrive at large-scale scanning facilities, where their bindings are removed so that pages can be processed at great speed.
Yet Maiberg’s analysis reaches far beyond the striking dystopian image of books being destroyed. What happens when knowledge changes form, changing ownership and accessibility? A book can circulate between readers, libraries and generations. Once destroyed and absorbed into a proprietary AI system, however, its contents survive virtually, encoded within a model whose outputs and rules of access are controlled by a private company. The paradox is striking: an industry seeking to ingest ever more human knowledge may simultaneously will erase the very sources of that knowledge.
Maiberg is careful not to portray this process as a diabolical campaign to rid society of books. His more unsettling interpretation is that the destruction is driven by something more mundane: indifference. In the race to build increasingly capable AI systems, books become inputs to be optimized, and their destruction becomes an acceptable operational cost. That logic, he argues, belongs to a broader concentration of informational power already visible in search engines and social media. And now AI is taking it to a whole new level.
The investigation therefore raises a larger question in the world of culture and academia: as more of the world's accumulated knowledge is mediated through AI systems, who controls the infrastructure through which that knowledge can be accessed?
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.