Skip to main content
Models & Technology

Rare Books Shredded for AI Training Spark Outrage; Book Database ISBNdb Takes Down AI-Related Test Page

To obtain high-quality training data, AI companies are purchasing large quantities of physical books and destroying them, a practice that has sparked widespread controversy. Book suppliers have received enormous orders, leading industry insiders to speculate that AI companies are the buyers behind them, while the database website ISBNdb has been accused of acting as an intermediary. Although the company has deleted the relevant promotional content and denied the claims, similar projects such as “Project Panama” were exposed long ago. The depletion of #AI training data# sources has driven companies to resort to the extreme measure of using physical books. #AI large-model training#

Rare Books Shredded for AI Training Spark Outrage; Book Database ISBNdb Takes Down AI-Related Test Page

This week, news that large quantities of valuable books were being shredded solely to provide data for training large language models (LLMs) attracted widespread attention from book lovers and ordinary internet users.

Rare Books Shredded for AI Training Spark Outrage; Book Database ISBNdb Takes Down AI-Related Test Page

According to the reports, some book suppliers recently received unusually large book orders. Industry insiders speculate that AI companies want to obtain higher-quality, cleaner training data while minimizing copyright risks as much as possible. The focus of the controversy, however, has turned to the book database website ISBNdb. Many believe that the site not only encouraged this practice but even helped connect AI companies with book suppliers. As public criticism quickly intensified, ISBNdb began trying to distance itself from the controversy and deleted the relevant content it had previously published.

In its latest statement, ISBNdb said: “We have taken note of recent reports about a marketing landing page on our website, and we understand the concerns it has raised. We do not train AI models, nor have we ever engaged in related business. That page was merely a test of market demand; the related service was never actually launched, and we have now removed the page.”

A few days earlier, 404 Media reported that multiple book suppliers had suddenly received large numbers of book-purchasing orders. As schools and libraries have faced tight budgets and declining procurement needs in recent years, many suppliers have been under considerable operating pressure and were therefore more likely to accept such large orders.

At the time, many people suspected that AI companies were the buyers behind the orders. ISBNdb became the focus of public attention because it had launched a service that appeared to be designed specifically to help AI companies establish connections with book-storage organizations.

The promotional page, which has since been deleted, once stated: “The world’s highest-quality AI training data is sitting on bookshelves. Books contain human knowledge that has been curated, peer-reviewed, and refined through professional expertise. Their level of structure cannot be replicated by any web crawler; their content is dense, edited, and authoritative.”

In fact, this is not the first time such a practice has been exposed. As early as January this year, a publicly available court document revealed Anthropic’s internally codenamed “Project Panama.” The project aimed to purchase as many books as possible, scan them digitally, and then destroy the physical copies.

Anthropic later reached a settlement with the authors totaling 1.5 billion US dollars (note: approximately 10.147 billion yuan at the current exchange rate). However, court documents made public afterward indicated that the practice itself was considered legal; what truly concerned the company was the negative public reaction it could generate. Anthropic was therefore willing to pay a high price to prevent the practice from being exposed.

Now that these details have become public, the large number of recent orders for valuable books has also come under closer scrutiny.

An anonymous bookseller told 404 Media: “This helps me clear out inventory that would otherwise be very difficult to sell, and it also allows me to make money. On the other hand, I don’t like what these books will ultimately be used for, and I don’t want to see some rare books turned directly into pulp.”

The report said that as AI companies find it increasingly difficult to obtain “clean” training data, they are looking for new data sources.

The vast amount of low-quality content on the internet, along with the growing amount of content generated by AI itself, could cause a “negative feedback loop” in model training. As a result, companies want to obtain higher-quality data while avoiding public attention as much as possible.

This summer, film company A24 announced a partnership with Google to provide training materials for DeepMind, in the hope of helping its media-generation models break free from the large volume of low-quality, stylistically inconsistent AI content.

As for why physical books must be destroyed, the report argues that it is neither to conceal evidence nor to preserve rare knowledge.

Although the documents related to “Project Panama” did not explain the specific reason for destroying the books, the most likely explanation remains cost reduction. In reality, digitally scanning books does not necessarily require damaging the originals. However, if the entire process prioritizes low cost and high efficiency, the bindings may be removed directly to obtain flatter, clearer scans, after which the damaged books can be disposed of as waste paper.