Photo: MYSURU, INDIA - AUGUST 05: Syed Khayum, a volunteer arranges books in the Anke Gowda Jnana Prathistana (Knowledge Shrine) on August 05, 2025 in Kennalu village, near the southern city of Mysuru, India. The Knowledge Shrine, founded and managed by 76-year-old Anke Gowda, a former sugar factory worker with no formal training in library science, is home to one of the world’s largest private book collections. Housing over 1.5 million titles across 22 Indian languages and several foreign languages, the collection spans literature, history, science, and the arts—all of which have been collected over decades with the intention of making learning accessible to the local community and to the citizenry in the form of a public library. However, the local community’s lack of interest in the library along with the absence of proper cataloguing and the growing burden of daily maintenance pose significant challenges to its preservation. (Photo by Abhishek Chinnappa/Getty Images)
The artificial intelligence (AI) industry is already under fire for building massive data centers next to quiet neighborhoods, sucking up vast amounts of water, and straining power grids. But now it may have a new public relations crisis on its hands – backlash over the destruction of millions of rare books.
According to recent reporting, Anthropic, the company behind the Claude chatbot, has bought up millions of books – including thousands of rare books – to train its AI model. To more efficiently scan the books into its servers, the company slices off their bindings and then recycles the used pages.
Anthropic launched the book-chopping endeavor in early 2024 under the name Project Panama. Internal documents and court records reviewed by The Washington Post reveal the scale of the closely guarded operation.
One internal planning document reportedly described Project Panama as an “effort to destructively scan all the books in the world.” Anthropic stored the resulting files in a searchable internal library. Other project documents indicated that the company allegedly did not want the operation publicly known.
Anthropic says it bought ordinary used books through mainstream retailers and did not target rare or antiquarian editions. But Australian booksellers have reported unusual orders for obscure and specialized titles, raising fears that Anthropic and other AI companies may be targeting rare books in order to monopolize access to the information in them.
One shipment from a bookseller included a 1970s soil-mechanics manual, a book on New Zealand cycling, an anthology of early Australian poetry, and a local history of a Melbourne suburb. The poetry collection had sat on the seller’s shelf for 20 years.
The unusual orders do not prove that every book is destined for an AI scanner. At least one company that has purchased large quantities of rare books, however, openly bragged about anonymity as part of a book-sourcing service for AI developers.
ISBNdb, a commercial book database, advertised a proposed service that would source as many as one million printed books for AI developers. The company promised strict nondisclosure agreements and pitched lawfully purchased books as an alternative to legally risky digital copies. Using digital copies found online could expose AI companies to copyright claims if those files were uploaded or distributed without the publisher’s permission, while purchasing a physical copy gives the company a clearer legal basis for scanning it.
After the offer drew scrutiny, ISBNdb removed the pages and said it had only been testing customer interest and had never launched the service.
Some groups have challenged Anthropic’s actions in court, but with little success. In Bartz v. Anthropic, U.S. District Judge William Alsup ruled that Anthropic’s one-for-one conversion of lawfully purchased print books into digital files for its internal library qualified as fair use. Because Anthropic destroyed each paper copy after scanning it, the court treated the process as replacing one format with another rather than creating an additional copy.
This destruction is not technologically necessary. Tesla and SpaceX founder Elon Musk said he instructed the SpaceX AI team to preserve rare books in a library and “scan them the hard way” rather than cut off their spines. The slower method may cost more, but it shows that AI companies do not have to choose between improving their models and preserving the original copies.
Destroying one copy of a widely available bestseller is one thing. But destroying an obscure engineering manual or local history is another matter entirely. When few copies remain, dismantling even one can reduce public access to information, images, or firsthand accounts of historical events that may not exist online.
Even a low shelf price does not make a rare book’s contents worthless. In fact, its obscurity may make it especially valuable to AI developers because it is less likely to appear in existing datasets. As machine-generated text spreads across the internet, older books offer a distinctly human source of training material.
But feeding a book into an AI model does not create a public digital edition. Readers cannot reliably ask a chatbot to reproduce the book, inspect every page, or identify the passage supporting an answer. The company controls the interface through which the public encounters the material. It can alter it, restrict access, or withhold it altogether.
Plus, an AI-generated answer about a book is not the same as access to the book itself. A chatbot may describe the source accurately, but it may also omit context, combine separate ideas, or confidently invent false details. Readers can catch those errors only when they can verify an AI answer against the original source.
A physical book provides an independent record. It preserves page numbers, illustrations, footnotes, and surrounding context. It also distributes custody of knowledge. Dispersed copies can survive in homes, used bookstores, universities, and public libraries, where unrelated readers can examine the same source and challenge one another’s interpretations.
AI companies should be free to learn from books, but they should not destroy scarce originals when libraries, archives, and nondestructive scanning methods offer another way. Technological progress should expand access to human knowledge, not concentrate it in the hands of a few AI giants.
Comments