Why Anthropic is destroying books — Kathryn James traced it
Kathryn James follows Project Panama from warehouse cutters to court, explaining why clean human prose became premium AI material.
Anthropic bought usable books, removed their spines, scanned them and threw them away. Vandalism usually requires less paperwork.
The codename: Project Panama.
An internal memo surfaced in Bartz v. Anthropic, examined by Kathryn James in The Guardian on August 5, 2026. Its mission:
Project Panama is our effort to destructively scan all the books in the world.
The memo explained the codename:
Why use a codename? … [B]ecause we don’t want it to be known that we are working on this.
A company that creates a codename, hires logistics staff, rents a warehouse, compares vendors and develops a legal theory has a strategy.
Why destroy books? Kathryn James found a boring answer with terrifying consequences: scanner economics and American copyright doctrine favored the same hydraulic cutter.
Project Panama was an industrial operation
According to James’s review of court records, Anthropic hired an experienced logistics manager, sourced books, staffed a warehouse and employed specialist digitization contractors.
Court exhibits showed labelled books on organized shelves, with employees, stacks and carts. Someone planned intake, scanning throughput and disposal.
Warehouses conceal endless decisions: equipment approvals, vendor contracts and the cheap scanner that jams every 300 pages, ruining Tuesday.
Judge William Alsup described the vendors’ work:
stripped the books from their bindings, cut their pages to size, and scanned the books into digital form – discarding the paper originals
Loose pages race through document scanners. Bound books need slower overhead cameras, flatbeds or V-shaped systems that support the binding while photographing each page.
Project Panama chose speed.
I understand the calculation. A startup buying millions of books models scanning costs, failures, staff and warehouse time. Unless an executive protects preservation, it becomes an expensive spreadsheet row.
Secrecy is harder to defend. Anthropic spent heavily knowing footage of pallets entering a spine cutter would look dreadful.
The memo said all the books in the world. I grew up in Ivrea, Olivetti’s town, and admire grand engineering ambitions. Even by Olivetti standards, this required espresso and a reread.
Book-destruction supply chains are deliberate: procurement compares bids, lawyers assess exposure, and executives decide negotiating with authors creates more friction than buying and cutting used copies.
That gets ugly.
AI slop made old books more valuable
Anthropic wanted books for edited prose, sustained arguments, unusual language and stories surviving beyond six seconds of TikTok attention.
The court materials described the prize:
well-curated facts, well-organized analyses, and captivating fictional narratives
Anthropic hoped books would help Claude match human authors’ accuracy and appeal. After a decade of push notifications, Slack reactions and refrigerator-reorganization videos, AI rediscovered books.
Bravissimo.
The pre-2022 cutoff matters. Books published before generative AI exploded probably lack ChatGPT, Claude or Gemini prose. Paper is now a rough certificate of human origin.
That certificate gains value as synthetic text fills the web. Models generate articles, product descriptions and forum answers; later models scrape them for training. This repetition can cause model collapse as systems learn increasingly distorted synthetic distributions.
TechRadar’s July 2026 analysis linked pre-2022 demand directly to this risk. Tom’s Hardware reported that professionally edited print offers cleaner human-authored material than much of today’s web.
The feedback loop is magnificently stupid. AI companies polluted online spaces with cheap prose, making old human writing scarce enough to extract from secondhand bookstores.
Molto Silicon Valley.
I first dismissed this as paper sentimentality, like keeping my stained Marcella Hazan cookbook when its recipes exist online.
I was wrong.
Clean human language now helps determine who builds the strongest models. Anthropic’s private corpus can improve Claude while excluding researchers, libraries and smaller European AI companies. The book leaves circulation; its value stays inside one American corporation.
Anyone who called long-form writing obsolete should notice that some of Earth’s richest companies spend millions ingesting books because sustained human thought is premium raw material.
Copyright law rewarded the cutter
Judge Alsup concluded that Anthropic could buy a book, make a digital copy for internal model development and destroy the paper. Under these facts, the private scan replaced the purchase.
He summarized:
One replaced the other.
Alsup stressed that the collection remained closed:
There is no evidence that the new, digital copy was shown, shared, or sold outside the company.
Destruction strengthened Anthropic’s argument. Keeping the book and scan created two usable copies; discarding the paper framed it as format conversion: one object entered, one private file survived.
American law does not require destroying every scanned book. Alsup assessed Anthropic’s process in a district court ruling, not binding nationwide appellate precedent. Another judge may disagree.
Corporate incentives rarely await a Supreme Court FAQ. Lawyers already have a favorable pathway.
One used book, one internal scan and no original can strengthen fair use while cutting costs. Copyright law gave the cutter legal value.
Google Books generally borrowed library books, photographed them non-destructively and returned them. Tom Turvey, formerly involved in Google Books partnerships, later joined Project Panama, according to court reporting summarized by GIGAZINE and Tom’s Hardware.
Anthropic bought used books, destroyed them and kept the corpus private. The warehouse received waste paper; Anthropic received proprietary training data.
Once a court validates the cheaper workflow, finance templates it, vendors package it and competitors copy it while communications teams call it a “digitization initiative.”

Project Panama converted a purchased physical book into a private digital file. The scan survived. The book did not.
Image alt text: Why Anthropic is destroying books through Project Panama’s destructive scanning process.
The $1.5 billion settlement taught one lesson
The settlement and destructive-scanning ruling concern separate book pools.
Anthropic accumulated more than 7 million pirated books. Authors Andrea Bartz, Charles Graeber and Kirk Wallace Johnson sued in 2024, arguing that pirate libraries supplied their work for Claude AI training.
The court distinguished those files from purchased scans. Training and converting legally bought books received favorable fair-use treatment; downloading and retaining pirated copies did not.
Anthropic agreed to a $1.5 billion settlement in 2025. Final approval came July 27, 2026, covering approximately 500,000 works at roughly $3,000 per eligible title. PC Gamer reported claims for 91% of eligible works.
A huge payment, but a Post-it-sized lesson for AI labs: get a receipt first.
Kathryn James reports that Anthropic CEO Dario Amodei considered copyright clearance a drawn-out legal and commercial slog. Negotiating across thousands of publishers, estates and territories does sound like eternity in an airport lounge with DocuSign.
Creators face the reverse: their books improve a commercial model, while only the used-book seller gets paid.
The settlement compensates piracy claims. It does not generally require AI companies to negotiate with authors before training on lawfully purchased copies.
Creative Bloq’s analysis drew the distinction sharply: the case punished how Anthropic assembled part of its library while changing much less about its use of legally acquired books.
Because the case settled, no appellate court bound future AI copyright cases. Google, Meta and OpenAI face separate disputes with different facts.
Authors won money. AI labs got a procurement lesson.
Strange buyers are clearing old shelves
Project Panama is documented. A murkier trend brought Australian and European booksellers large, apparently random orders for obscure titles.
Guardian Australia reported on August 2, 2026, that Delfina Manor of Good Reading Secondhand Books in Benalla, Victoria, received a Zoom Books order for about 30 to 40 books filling three boxes.
Manor described the work:
They ordered, paid in advance, and they didn’t quibble over the postage … [but] it would have been about 30 to 40 books, and so finding them, packing them, making sure you hadn’t missed one out … It was just driving me nuts,
John Sainsbury of Sainsburys Books in Melbourne reported similarly random, price-insensitive demand for niche, decades-old stock. One order combined a 1970s soil-mechanics manual, Born to Thunder: Champions of New Zealand Cycling, a 1982 Early Australian Poetry collection and a Hawthorn local history.
I want to meet the reader planning that weekend.
The poetry book sold for $9 after roughly 20 years on its shelf. A bookseller moved dead stock, but nobody knew whether it went to a reader, arbitrage warehouse or industrial cutter.
Nick and Jenny Dawes estimate Grant’s Bookshop holds around 500,000 books in its warehouses. Nick told The Guardian he would feel conflicted if someone bought everything for cutting and would refuse to destroy the only known copy.
Fair enough. Used-book stores need revenue; a dusty 1987 engineering manual does not pay rent.
Uncertainty creates the risk. Anthropic says it has never bought from Canadian reseller Zoom Books, and its programs neither buy nor destroy rare or antiquarian titles. Zoom Books says it resells books intact and does not digitize them, but withheld customer identities under confidential commercial agreements.
I cannot responsibly connect the Australian orders to Anthropic. The evidence does not support it.
ISBNdb adds another mystery. Archived marketing reported by 404 Media and PC Gamer offered printed-book sourcing for LLM training, from 1,000 to 1 million books, promising:
your identity, strategy, and acquisition targets are never disclosed.
The archived copy understood the optics:
‘AI company destroys two million books’ is not a headline that generates sympathy.
ISBNdb removed the offer, saying it only tested demand and never purchased, scanned or sold a physical book.
An anonymous specialist bookseller told 404 Media that weekly sales rose from around 20 books to hundreds after the surge. The orders shared little but ISBNs, suggesting lists built from bibliographic databases.
No Library of Alexandria cosplay is needed. Opaque buyers can quietly remove uncommon books while sellers cannot assess preservation risk. Confidential targets make stewardship nearly impossible.
A corporate corpus makes a lousy library
Anthropic proposes a “forever” research library. Bold. My self-hosted Docker stack develops opinions after routine updates, so I distrust corporate eternity.
Preservation libraries document provenance, protect exceptional copies and provide credible future access. Corporate training corpora guarantee none of that.
The public lacks a complete inventory of Anthropic’s destroyed titles. Researchers cannot freely inspect scans; historians cannot examine originals.
OCR captures text. It misses plenty.
Books carry marginalia, ownership stamps, printing errors and handwritten recipes. Paper reveals production methods; bindings distinguish cheap from elite editions; inscriptions connect objects to families, institutions or political movements.
I’m Italian, so technology arguments must become lunch. A stained community cookbook with a handwritten lard substitution records how a family cooked. Clean OCR preserves the official recipe but deletes the improvisation.
Tim White of Melbourne’s Books for Cooks collects evidence of where food meets society. He told Guardian Australia that rare booksellers handle objects with stories.
His verdict:
If it’s a book that has that storytelling element to it, or it’s a one-off, it’s horrific.
Rare bookseller Gwenyth Todd screens buyers because cutting books apart to sell illustration plates physically appalls her. Profitable destruction predates AI; language-model companies can industrialize it beyond plate dealers’ dreams.
Compared with protections for other artifacts, the United States gives books few cultural-heritage safeguards. No practical endangered-species list protects a regional history with three known copies.
Any AI lab buying books industrially should follow six rules:
- Scan scarce, annotated, antiquarian and out-of-print works without cutting their bindings.
- Check library catalogues and bibliographic databases for rarity before destruction.
- Publish an inventory of every title destructively scanned.
- Deposit preservation-grade files with a trusted library under controlled access where copyright requires it.
- Offer authors and publishers direct licensing routes.
- Allow independent audits of acquisition vendors and scanning facilities.
These rules cost money. Good. A billion-dollar model built from humanity’s written record can afford a rarity check.
On July 27, 2026, Elon Musk said he instructed the SpaceXAI team to preserve rare books and scan them “the hard way.” That is a useful minimum commitment. Patron saint of librarians remains premature.
The equipment exists. Google scanned non-destructively; the Internet Archive uses preservation-oriented systems. Datamation Information Services, linked to Project Panama through court reporting, offers destructive and non-destructive methods.
The cutter is a business choice.
Earth’s most advanced language machines scour used bookstores because they cannot manufacture what they need: a deep record of human thought from before the machines arrived.
Books were supposedly dead, writers replaceable and every meaningful idea feed-sized. Now pre-2022 human writing is valuable enough to warehouse and scan into billion-dollar products.
Before destroying a book, AI companies should answer two questions: Is this copy replaceable? Who can access what survives?
By 2028, I expect major AI procurement contracts to require rarity screening and public title inventories, through regulation or publishers withholding cooperation. Otherwise, our descendants may find humanity’s written record preserved in proprietary model weights while readable copies entered recycling bins.
Frequently asked questions
Why is Anthropic destroying books?
Anthropic destroyed purchased books because removing their bindings enabled faster, cheaper high-speed scanning. Destroying each paper copy also supported its argument that one purchased object had been converted into one private digital file rather than duplicated, strengthening the fair-use position accepted by Judge William Alsup.
Did the Anthropic copyright settlement cover books it legally purchased and scanned?
The $1.5 billion settlement covered claims involving pirated books, not the separate pool of physical books Anthropic legally purchased and scanned. The court treated downloading and retaining pirated files differently from converting purchased books into private digital copies for internal model development.
Why are pre-2022 books valuable for AI training?
Pre-2022 books offer edited, sustained human writing that is unlikely to contain text generated by modern systems such as ChatGPT, Claude or Gemini. As synthetic prose spreads across the web, older printed books provide cleaner human-authored material and reduce exposure to recursive training on AI-generated content.
Sources
- Why is Anthropic destroying books? | Kathryn James
- ‘More than just objects’: Australian booksellers raise alarm over ‘horrific’ destruction of rare titles to feed AI
- Company Offering Printed Books to Train AI Stops After 404 Media Coverage
- AI companies are anonymously buying and destroying millions of books through middleman services to avoid headlines about AI companies buying and destroying millions of books
- Company that said it could scan and destroy books for AI data-harvesting has deleted that part of its website: 'no such service was ever brought to life'
- AI companies are reportedly shredding millions of books after using them to train AI models — tech giants outsource to middlemen to secretly buy up books for training material