BlogIndustry Analysis

AI Companies Are Shredding Rare Books to Train Their Models. A Judge Said It's Legal.

AI labs are bulk-buying rare books, slicing off the spines, scanning the pages, and shredding the originals. ISBNdb facilitates orders of up to a million books with anonymous buyers and NDAs. Pre-2022 books command a premium because they're "structurally clean" of AI text. A federal judge called it fair use. The books are gone forever.

Chethan·July 27, 2026

AI companies have found a new way to get training data. They're buying rare books, slicing the spines off with hydraulic cutters, scanning the pages through industrial machines, and shredding what's left. A service called ISBNdb facilitates orders of up to a million books at a time and keeps the buyers anonymous. Pre-2022 books fetch a premium because they're guaranteed free of AI-generated text.

This is happening right now. A federal judge already said it's legal. And the books being destroyed aren't just paperbacks — they include rare, out-of-print volumes with only a handful of surviving copies anywhere in the world.

Let me explain what's actually going on, because the details are worse than the headline.

The Pipeline

Here's how it works. AI labs need training data. The internet — their traditional feeding ground — is increasingly contaminated with AI-generated text. Every blog post written by ChatGPT, every auto-generated SEO page, every synthetic Reddit comment is poison for a model being trained to understand human writing. Researchers have a name for this: model collapse. Train an AI on AI output, and it degrades. Fast.

So the AI labs went looking for data that's provably human. And they found it sitting on bookshelves.

Print books published before 2022 are what the industry now calls "structurally clean." No AI touched them. They're pure human signal. ISBNdb, the company brokering these bulk purchases, puts it plainly on their own website: "The world's best AI training data is sitting on a shelf." They call pre-LLM-era books "structurally guaranteed to be free of this contamination." They even refer to AI-generated content as "modern poisoning tools." These are the people selling the books, and they're calling their own product's competitors poison.

The process itself is industrial. Buyers purchase books in bulk — thousands, sometimes tens of thousands at a time. The books go through destructive scanning: a hydraulic blade slices off the spine, the loose pages feed through high-speed imaging equipment, and the digital text gets extracted for model training. The physical book is destroyed in the process. Not stored. Not archived. Destroyed.

ISBNdb facilitates orders from 1,000 to one million books. They keep buyers anonymous. They offer NDAs as a feature — literally a selling point on their website. They coach clients to call it "digital preservation." And they're remarkably self-aware about the optics, with their own site acknowledging: "'AI company destroys two million books' is not a headline that generates sympathy."

They built the business anyway.

Anthropic and Project Panama

The most documented case comes from Anthropic. Court records from a recently settled copyright case revealed that the company bought large numbers of printed books from second-hand sellers, removed the bindings with a hydraulic cutting machine, and scanned the pages using industrial imaging equipment.

Anthropic had an internal name for this: Project Panama. According to court documents, it was a plan with a straightforward goal — destructively scan every book in the world. The company reportedly spent tens of millions on it and hired the former head of Google Books partnerships to lead the effort, tasking him with obtaining "all the books in the world."

To be clear about the scale: this isn't a side project. It's a data acquisition strategy with a budget, a team, and a target of comprehensive coverage.

The Judge Said It's Fine

Here's where it gets legally interesting. When this ended up in court, the argument centered on the first-sale doctrine — the principle that when you buy a physical book, you own that physical copy and can do what you want with it. Rip it up. Burn it. Or scan it and feed the text to a language model.

A federal judge found the scanning process to be "transformative," which made it eligible for fair use protection under US copyright law. The reasoning: because the physical book was destroyed during digitization, only one copy existed at any given time. The digitized version replaced the physical one rather than duplicating it.

This is a genuinely novel legal argument, and it's now one of the defining decisions shaping how AI companies acquire training material. If you can scan a book and destroy the original, and that counts as fair use because no copy was made, then every physical book in the world is legally scannable — as long as you shred it after.

Anthropic did face separate copyright trouble. In June 2025, the company agreed to a settlement valued at up to $1.5 billion in litigation over the unauthorized use of pirated digital books. But that settlement was about digital piracy — downloading copies of books they didn't own. The fair use ruling on physically purchased, destructively scanned books remains untouched.

The distinction matters. Pirating a digital copy: illegal, $1.5 billion penalty. Buying a physical book, scanning it, destroying it: fair use, no penalty. The legal system has created a clear incentive structure, and it points in one direction. Buy physical. Shred physical. Move on.

What's Being Lost

A bookseller told 404 Media that his sales jumped from fewer than 20 books per week to several hundred almost overnight in April. The buyers didn't seem to care about genre, topic, or condition — only that each book had an ISBN. That detail alone tells you everything. They weren't curating a collection. They were harvesting raw material.

His inventory includes rare, foreign-language, and out-of-print titles. Books that might be among the last few surviving copies in commercial circulation. And they're going into a shredder.

"I personally have mixed feelings about all of this," he said. "It benefits me financially... On the other hand, I don't like the end-use, and I don't like that uncommon books are being pulped."

This isn't happening only in the United States. Rare booksellers in the Netherlands have reported similar unexplained bulk purchases, with buyers who won't identify themselves and orders that appear out of nowhere.

Here's what makes this different from every other AI data controversy. When AI companies scraped the internet, you could re-upload. When they torrented library collections, the originals stayed on library shelves. When they copied code from GitHub, the repos were still there.

Destruction is irreversible. Once the last physical copy of an 18th-century botanical text goes through a shredder, it's gone. The digital scan captures the text. It doesn't capture the binding, the marginalia, the printing techniques, the paper composition, the provenance. A scan is not a book. It's a transcript.

And here's the thing that should genuinely unsettle you: storing a book after scanning costs money. Warehouse space, climate control, inventory management. Shredding it costs nothing. So guess which option the market chose.

The Data Economy's Dirty Secret

This story reveals something structural about the AI industry that doesn't get talked about enough. The entire AI boom is built on a data extraction model that treats human-created content as a free natural resource to be mined. When the easy mines — the public internet — started running dry, the industry didn't slow down. It found harder-to-reach deposits and built the infrastructure to extract them.

Physical books are the new oil fields. And like oil extraction, the process is destructive, the local communities (in this case, booksellers and archivists) get minimal benefit, and the long-term environmental cost (cultural loss) is externalized.

ISBNdb's own marketing is accidentally revealing about the state of the industry. The fact that pre-2022 books command a premium is an admission that the internet is broken as a training data source. The web is now so saturated with AI-generated content that AI companies would rather shred rare books than risk training on it. They're destroying physical artifacts to escape the pollution their own products created.

There's a circularity here that would be funny if it weren't real. AI companies generated so much text that they ruined the internet as a training data source. So now they're destroying books to get clean data. The models they train on that clean data will generate more text, which will further contaminate the internet, which will increase demand for more physical books to shred. It's a feedback loop, and the physical record of human knowledge is the fuel being burned.

Why Nobody Can Stop It

The legal framework is settled, at least for now. Fair use. First sale doctrine. Transformative use. These are established principles that predate AI by decades, and they happen to create a perfect shield for destructive scanning.

Copyright law was designed to regulate copying. It says you can't make unauthorized copies of someone else's work. But the courts have decided that scanning a physical book you own and destroying the original isn't copying — it's transformation. The law has no mechanism to address the destruction of cultural artifacts for commercial purposes, because nobody imagined this scenario when the framework was written.

There's no "cultural preservation" exception to the first-sale doctrine. There's no requirement to archive a book before you shred it. There's no registry to track which books have been destroyed. The bookseller who sells a rare volume to an anonymous buyer has no way of knowing, and no legal right to ask, whether it's heading to a scanner.

The only thing that could slow this down is public pressure or new legislation. But the buyers are anonymous. The transactions are private. The NDAs are signed. And by the time anyone notices a book is gone, it's already been pulped.

The Real Cost

We talk a lot about AI safety in terms of what models might do — generate misinformation, automate weapons, replace jobs. Those are important conversations. But there's a quieter, more permanent kind of harm happening right now, in warehouses, with hydraulic cutters and industrial scanners.

An 18th-century botanical text that survived two and a half centuries of human history — wars, floods, moves, estate sales, the slow attrition of time — doesn't get to decide whether it becomes training data for a language model. It just gets bought, scanned, and shredded. The digital text lives on inside a model's weights, transformed into statistical patterns that nobody will ever read. The book itself becomes pulp.

Brian Roemmele, who has been tracking this issue for years, called it "The Great Forgetting." That's not hyperbole. Every book destroyed is a small piece of the human record converted into something unrecognizable — a set of floating-point numbers in a model that can approximately reproduce the patterns but has no relationship to the original artifact.

The scan drops what the physical book carries. The marginalia — handwritten notes from previous owners. The printing history — which edition, which press, which corrections. The binding — how it was made, what materials were used, what it tells us about the era. All of it, gone. What remains is text stripped of context, fed into a system whose entire purpose is to forget the specifics and remember the patterns.

This is the trade the AI industry has made, and a federal judge signed off on it. Your cultural heritage, converted to model weights, physical originals destroyed. Legally. Quietly. At industrial scale.


If this bothers you — and honestly, it should — the question is what to do about it. Supporting independent booksellers and archives matters. So does paying attention to the data practices of the AI tools you use every day. Not every AI company is shredding books. But the ones building the biggest models need the most data, and the clean data is running out.

At CopperRiver, we run open-source models on your own machine. Your data stays yours. No book shredding required.

#AI training data#fair use#copyright#cultural preservation#Anthropic

Try CopperRiver yourself

A desktop AI assistant that browses, codes, and automates. Plans from $9/mo.

Read next

AI Companies Are Shredding Rare Books for Training Data | CopperRiver