Across Europe and North America, antiquarian booksellers are reporting an unusual new type of customer: well‑funded, anonymous buyers willing to pay several times market price for obscure, often out‑of‑print volumes — on the condition that the books can be taken apart, scanned, and never resold. The surge, first widely reported in connection with Anthropic’s “Project Panama,” reflects a structural shift in how large AI labs source text: away from the open web and toward physical books that are both legally purchasable and “free of AI slop.”
AI Companies Buying Tons of Old Books for “Slop‑Free” Training Data
Investigations by 404 Media, Boing Boing and other outlets reveal that a book‑data company called ISBNdb is now brokering bulk purchases of printed books — from 1,000 up to 1 million volumes per order — specifically for AI labs. ISBNdb markets pre‑2022 print books as “structurally guaranteed” to be free of AI‑generated contamination, positioning them as high‑quality training data compared with a web increasingly clogged with synthetic text.
The company’s sales copy, quoted by 404 Media, frames print collections as “curated, peer‑reviewed, domain‑specific human knowledge” that web crawls cannot replicate, and openly acknowledges an “optics problem” — admitting that “AI company destroys two million books” is not a headline that generates public sympathy. Booksellers interviewed by 404 Media and Boing Boing say they’ve seen order sizes and purchasing patterns shift dramatically since spring 2026, with buyers showing “total disregard for price,” a tell‑tale sign of deep‑pocketed AI clients.
Destructive Scanning: Cut the Spine, Save the Data, Shred the Book
The process used in many of these operations is what librarians call “destructive scanning.” To maximise throughput and image quality, buyers slice off a book’s spine, separate the pages, and feed them through industrial‑grade imaging machines, typically at 600 DPI or higher. After scanning, the paper originals are pulped or shredded; their only surviving form is as digital files and, ultimately, as weight parameters inside large language models.
Court documents in the US Anthropic litigation show that the company “cut up millions of printed books, scanned them, and used them solely for AI training, then discarded the originals,” a practice Judge William Alsup later held to be “clearly transformative” fair use for the purpose of model training. Unlike non‑destructive preservation projects such as Google Books or the Internet Archive’s scans, this model of destructive scanning guarantees that the physical copy is removed from circulation after digitisation.
Rare‑Book Dealers Sound the Alarm
NL Times and other European outlets report mounting concern among antiquarian and rare‑book dealers that AI companies are systematically removing culturally significant volumes from the market. Dealers in Amsterdam and Leiden describe buyers, later linked via logistics and payment trails to major AI labs, paying 3–5× the usual price for obscure 16th–18th‑century editions, including early Dutch legal codes and Latin theological works that exist in very few surviving copies.
These transactions often come with unusual conditions: “no resale,” provision of high‑resolution single‑page scans, and confirmation that the physical artefacts will be destroyed after imaging, sometimes following secure‑destruction standards used for sensitive records. For the dealers, the immediate upside is a spike in sales; one described going from about 20 books a week to hundreds. But many worry that books which have survived wars, fires and centuries of handling are now disappearing quietly into shredders — their contents preserved, but their status as physical heritage erased.
Why AI Labs Want Pre‑2022 Print Books
Technical explanations for the shift are relatively clear. Researchers and practitioners have warned about “model collapse,” the degradation that can occur when models are trained repeatedly on AI‑generated output rather than human‑written text. As generative systems flood the web with synthetic content, finding large corpora of reliably human‑authored material becomes harder, especially outside mainstream languages and genres.
Pre‑2022 print books offer several advantages:
-
They are temporally guaranteed to predate large‑scale consumer generative AI.
-
They contain dense, edited prose across specialised domains — law, theology, science — that is underrepresented in publicly available digital corpora.
-
They can be lawfully purchased in bulk, with a paper trail that simplifies licensing arguments compared with web scraping.
Academic work on OCR‑augmented pretraining shows that high‑quality scans of historical, low‑resource languages can significantly improve model performance on specialised benchmarks, giving a clear performance incentive to ingest rare historical texts.
Copyright Law and “Receipt as License”: Why Destruction Can Be Rewarded
A recurring theme in analysis from outlets like CryptoSlate, AIWeekly and 404 Media is that existing copyright doctrine may be quietly rewarding destructive practices. In the Anthropic case, Judge Alsup reasoned that buying print copies, scanning them for internal AI training and destroying the originals — without redistribution of the text — counted as a transformative fair‑use application, particularly when the books were lawfully purchased.
Commentators summarise the logic as “the receipt is the new license”: once a company can show legitimate purchase of a physical book, it has a stronger fair‑use argument for creating internal digital copies for analytical uses such as model training, especially when no “competing” reading service is offered and the text is not made publicly available. In that framework, destruction of the physical copy can even help the fair‑use case by demonstrating that the buyer is not adding new copies to the market, but instead substituting the artefact with a non‑public digital derivative.
Critics argue this creates perverse incentives. Instead of digitisation and preservation, companies may be nudged toward one‑way consumption: buy, scan, destroy, keep the data. For rare works, that means access rights and preservation choices shift from public institutions to private labs optimising for cost, throughput and legal defensibility, not for posterity.
Anthropic’s Project Panama and the “Optics Problem”
Anthropic’s internal initiative sometimes referred to as “Project Panama” has become the emblem of the practice. Reporting and court filings show the company spent millions of dollars acquiring large numbers of printed books, cutting their spines and feeding pages through scanners as part of a concentrated effort to build training sets from non‑digital sources. While there is no direct evidence that Anthropic targeted first‑edition or uniquely valuable rare books, the scale of the operation has amplified wider concern that AI labs might not adequately differentiate between common paperbacks and scarce cultural artefacts.
ISBNdb, for its part, openly warns clients that “the optics problem is real” and promises strict NDAs for every engagement, a sign that the companies involved expect public backlash if specific deals become widely known. Industry blogs note that phrases like “strip‑mining the world’s old books” are already circulating in critical coverage, capturing anxiety that historical texts are being consumed as raw material rather than stewarded as heritage.
Cultural Heritage, Access and the Future of AI Training
What is not in dispute is that large numbers of books are being bought, scanned and destroyed to fuel AI training; the questions now revolve around which books and under what conditions. Rare‑book dealers stress that many of the titles being requested are obscure, non‑English works from the 16th–18th centuries, the kind of texts unlikely to be reprinted or digitised by mainstream publishers. If those volumes vanish from physical shelves, future scholars may have to rely entirely on corporate‑controlled digital reproductions, assuming they remain accessible at all.
Policy analysts and copyright scholars suggest several potential responses:
-
Stronger heritage‑protection rules for certain classes of rare or unique books, akin to export controls on cultural artefacts.
-
Clearer licensing frameworks for non‑destructive scanning partnerships between AI firms and libraries, balancing training needs with preservation.
-
Re‑examining fair‑use doctrine when destructive practices become widespread and materially affect public access to works that were previously obtainable in physical form.
For now, the practice is largely governed by market dynamics and court precedents that view internal AI training as a transformative analytical use. As AI labs push for ever‑larger, higher‑quality corpora, the tension between “clean data” and cultural stewardship is likely to intensify.