AI 6 min read

Cut the Spine, Feed the Model: The Quiet Destruction Behind AI Training Data

To scan a book quickly, you have to ruin it. The spine gets sliced off, the pages become loose sheets, and the sheets get fed through a high-speed document scanner. Then the paper goes in the trash. This is happening at industrial scale to build AI training corpora, and archivists and librarians are quietly furious about it.

One caveat up front: this isn’t a story trending on Hacker News this week. It surfaced in litigation documents and has been simmering in library circles for a couple of years without ever quite catching fire. So treat what follows as a map of the argument, not a readout of live outrage. The argument hasn’t gotten any less sharp for sitting still.

Why Not Just Use a Non-Destructive Scanner

The obvious objection first: overhead scanners exist. Robotic page-turners exist. Why cut anything?

Speed and cost. Photographing a bound book means someone — a human or a robot arm — turns every page. Text curves into the gutter near the spine, so you need software to flatten it back out. Errors compound. Throughput is maybe a few hundred pages an hour on good equipment.

Cut the binding off and the book becomes a stack of paper. A commodity sheet-fed scanner eats hundreds of pages per minute with no correction pass. When your target is tens of thousands of volumes, that gap is the entire project. It’s the difference between a pipeline and a hobby.

Destructive scanning isn’t new, and it isn’t uniquely an AI-industry sin. Anyone who has ever taken a stack of textbooks to a copy shop to get them turned into PDFs has done the same thing on a small scale. What changed is the volume.

The Paradox: Buying Books to Destroy Them Is the Safe Option

Here’s the genuinely strange part. Pirated copies of most of these books are already sitting on the open internet. Why go to the trouble of buying physical copies and shredding them?

Because of first sale doctrine. Under US copyright law, when you buy a legitimate copy of a book, that specific copy is yours to do whatever you want with — read it, resell it, burn it, guillotine it. And in 2025, a federal court held that buying print books, scanning them, and discarding the originals was fair use, precisely because it was a format shift of a copy the company legally owned. Downloading the same text from a pirate library was not fair use. Same words, same model, entirely different legal outcome.

Read that back and the incentive structure is almost comically perverse. Buy the paper and physically destroy it, and you’re on solid ground. Torrent the file and leave every physical book in the world intact, and you’re liable for damages. The law makes preserving the original the riskier path.

Nobody designed it that way. But that’s what the rules reward right now, and companies with legal departments respond to what the rules reward.

The Best Counterargument: They’re Just Used Paperbacks

The defense of this practice is stronger than critics usually admit. What gets cut up is overwhelmingly mass-market inventory — the kind of thing that shows up by the pallet at used book liquidators. Slicing the spine off a 2003 management title that had a print run of 200,000 does not erase a piece of human heritage. Arguably it rescues one: a book rotting in a warehouse becomes searchable text.

The problem isn’t the average book in the box. It’s that nobody looks in the box.

Bulk acquisition happens by weight and by pallet. There’s no triage step, because triage is exactly the labor the whole pipeline was designed to eliminate. Mixed into those pallets: out-of-print local histories, first editions from small presses that folded in the 1990s, signed copies, books with a previous owner’s marginalia running down every page. Librarians have been making this point for years, and it’s the one that actually lands — scarcity is usually discovered in retrospect. The 1972 zoning report nobody wanted becomes the only surviving record of a demolished neighborhood.

And a book is not only its text. Paper stock, printing method, binding style, bound-in advertisements, bookplates, library circulation stamps — a substantial share of what bibliographers actually study never makes it into a 300 DPI image of the text block. Scanning isn’t copying. It’s extraction. You take the layer you need and throw away the substrate.

The Real Loss Comes After the Scan

This story attracts exaggeration, so it’s worth being precise. There’s no evidence anyone is feeding medieval manuscripts or museum-grade rare books into a guillotine. The material being destroyed is, in the overwhelming majority of cases, ordinary commercial publishing. Aiming the outrage at the wrong target is how a legitimate concern gets dismissed.

The part that should bother you is what happens next. The scans don’t come back out.

Google Books, for all the litigation it generated, at least returned something to the commons — full-text search, snippet view, a public index of what exists. AI training scans return nothing. The physical original is gone, the digital surrogate sits on a private storage cluster, and the model trained on it is usually closed too. A public resource was consumed and converted entirely into a private asset.

That’s the actual grievance from the library world, and it’s more coherent than “they’re destroying books.” Destruction in service of preservation is a trade libraries have made for decades — microfilming programs did exactly this. The objection is to destruction with no preservation on the other side. Who pays and who collects.

What Would Actually Fix This

Three asks, all of them modest, none of them currently required by anything.

Triage before the blade. Establish criteria for pulling potentially rare items out of bulk lots before cutting. Publication date, print run, publisher status, physical annotations. Not perfect, but a rare-book flag on a checklist costs almost nothing next to the price of the books themselves.

Don’t bin the paper. Loose sheets still carry the paper stock, the printing, the marginalia. Boxed and shipped to an institution that wants them, a disbound book is a damaged artifact rather than a destroyed one.

Donate the images. Not the training corpus, not the model weights — just the page scans of books that no longer physically exist, deposited with a public archive. This is the smallest possible ask: leave a digital trace of the thing you removed from the world.

All three cost money. That’s the whole reason none of them happen.

The Missing Inventory

One book disappearing is nothing. Books disappear constantly — fires, floods, basement cleanouts, estate sales nobody catalogs.

What’s different here is the silence. Tens of thousands of volumes are being converted into text and stripped of everything else, with no record of what went in. If something irreplaceable was in one of those pallets, we will not find out, because there is no list. The loss and the record of the loss are destroyed in the same motion.

This was never a choice between advancing the technology and keeping the originals. Both are affordable. It’s just that right now, only one of them has anyone willing to pay for it.

AI copyright digital archives data ethics publishing

Comments

    Loading comments...