AI firms buy and destroy rare books to train LLMs

AI companies are buying up rare books to train large language models (LLMs) such as Claude, ChatGPT, and Gemini, and then destroying them, sparking anger among booksellers and collectors. A recent Wall…

AI companies are buying up rare books to train large language models (LLMs) such as Claude, ChatGPT, and Gemini, and then destroying them, sparking anger among booksellers and collectors. A recent Wall Street Journal report said sellers were astonished by a sudden rise in rare book sales to buyers with names like Red Sparrow Project and Blue Finch Project, who ordered hundreds of books at a time without haggling.

AI firms buy and destroy rare books to train LLMs

The books are taken to warehouses, where their spines are sliced, they are scanned, and then pulped. A U.S. court case involving Anthropic, creator of Claude, revealed Project Panama, an attempt to destructively scan all the books in the world. The court ruled that Anthropic's use of copyrighted books to train AI was not copyright infringement, as the digital copies were not shared or sold outside the company.

The incident raises questions about whether books should be treated as cultural objects or mere containers of text for AI. As reliance on LLMs grows, more such legal battles are expected over the destruction of irreplaceable historical documents.

Indian Opinion Analysis

This is not the first clash between AI training and copyright. In 2023, authors including John Grisham and George R.R. Martin sued OpenAI over unauthorised use of their works. The difference here is the physical destruction of artefacts: rare books are often unique, with marginalia, binding, or provenance that digital scans cannot capture. The Indian context matters too: the National Digital Library of India holds lakhs of rare texts, but there is no legal framework here governing whether an AI company can buy and destroy them. The Anthropic ruling is from a U.S. court and does not bind Indian courts, but it signals a permissive global trend. Watch for the next hearing in the U.S. class action against OpenAI, where a judge may decide whether training on copyrighted works is fair use.

The real question is who decides what gets preserved.


Source: thehindu.com

This brief was synthesised by AI from the source linked above.

Ask their opinion on this story
They have read this article, our coverage, and the web.
AI simulations of historical figures. Responses are generated from the historical record, not authentic statements.

0 Votes: 0 Upvotes, 0 Downvotes (0 Points)

Share your opinion

Loading Next Post...
Search Trending
Ask their opinion
Loading

Signing-in 3 seconds...

Signing-up 3 seconds...

All fields are required.