Skip to content
Home » Blog » Why AI Companies Are Destroying Books to Train Models

Why AI Companies Are Destroying Books to Train Models

  • by
Photographer: Iñaki del Olmo | Source: Unsplash

AI companies are destroying books because carefully edited books provide high-quality training material for large language models. In a widely reported case involving Anthropic, the company bought millions of print books, including antique books, removed their bindings, scanned the pages, and discarded the paper copies. The result was a searchable digital library that could be used to improve AI models such as Claude.

The important detail is that this story is not simply about "AI burning books." It is about a collision between three ideas: the commercial value of high-quality training data, the legal rules around copying and format shifting, and the cultural value of physical knowledge. A book can be successfully digitized while its original artifact is permanently lost, even when it is a decent print or a durable hardback copy.

That distinction matters to authors, publishers, librarians, historians, businesses, creators, and anyone who cares about who controls the information used to build artificial intelligence.

Why are AI companies buying and destroying books?

AI companies are seeking reliable, well-structured human writing for their training datasets. Printed books are attractive because they have usually passed through editing, fact-checking, publishing, and quality control. Compared with the constantly changing web, a book offers a stable and coherent source of language, facts, arguments, and storytelling.

There is also a growing concern about low-quality synthetic content online. When new AI systems are trained primarily on recent web pages, they may encounter more automatically generated articles, duplicated text, inaccurate summaries, and machine-written material. This can create a feedback loop in which one model learns from the mistakes of another.

Books offer a way to add older, human-created material to the mix. They contain subjects that may be missing from the modern web, including obscure history, specialist knowledge, regional perspectives, and writing styles that were developed before generative AI became widespread.

In short, books are valuable to AI companies because they are often:

  • Human-written and edited by writers
  • Long-form and information-dense
  • Stable enough to catalog and reuse
  • Available in large quantities through used-book sellers
  • Useful for teaching language, reasoning, explanation, and narrative structure
  • Produced with publication standards that can resemble bestseller-grade materials

The irony is difficult to miss. AI companies want the durability and accumulated wisdom of books, but their preferred scanning method can eliminate the physical object that carried that history.

How destructive book scanning works

Destructive scanning is a high-speed digitization process in which a book is cut apart so its pages can be scanned more efficiently. Instead of carefully opening a bound volume page by page, a machine removes the spine or binding, separates the pages, scans them in batches, and sends the remains for recycling or disposal. In some cases, this may be the only type available for scanning at an incredible scale, but it is also a disposable format for the physical object.

The process generally looks like this:

  1. Books are purchased in bulk from distributors, used-book sellers, or other suppliers.
  2. The books are sorted and cataloged using details such as title, author, subject, and ISBN.
  3. A cutting machine removes the binding and separates the pages.
  4. Production scanners capture the pages and create searchable digital files.
  5. Optical character recognition converts the page images into machine-readable text.
  6. The original paper copies are discarded or recycled.
  7. Selected digital copies may be cleaned, organized, and added to AI training datasets.

A June 23, 2025 order from the United States District Court for the Northern District of California described this process in the Anthropic case. The order said Anthropic purchased millions of print books, had service providers remove their bindings, scanned the books, and discarded the paper originals. The court record also described the company’s broader effort to build a central research library for future use.

The court document is important because it provides primary evidence rather than relying only on social media descriptions of the controversy. Read the Anthropic fair use order for the detailed account.

Photographer: Susan Q Yin | Source: Unsplash

Did Anthropic destroy millions of books?

Court filings and reporting indicate that Anthropic bought and destructively scanned millions of physical books as part of its Project Panama effort. The precise number of books scanned was not fully disclosed in the public filings, and some project documents discussed a potential volume ranging from 500,000 to two million books over six months.

That means it is important to be precise. The public record supports the claim that Anthropic acquired and destroyed very large numbers of print books. It does not prove that every rare or unique book was shredded, and it does not establish that all AI companies are using the same process.

The Washington Post reported that Anthropic spent tens of millions of dollars acquiring and cutting up millions of books before scanning the pages. The report also described internal plans to create a large, private library of books for AI research and model development. You can read the Washington Post investigation into Anthropic’s book-scanning project for additional reporting and context.

This distinction between documented facts and speculation is essential. The strongest criticism does not require exaggerated claims. The confirmed facts are already unsettling enough: a company can treat physical books as raw material, convert them into private data, and eliminate the originals as part of an efficient industrial workflow.

Is destroying a scanned book legal?

The short answer is that one court ruled that Anthropic’s purchase, digitization, and use of lawfully acquired books for AI training qualified as fair use in the circumstances before it. That ruling does not mean that every use of copyrighted books for AI training is automatically legal.

The June 2025 Anthropic order distinguished between two parts of the case:

  • Anthropic’s use of lawfully purchased books that were destructively scanned
  • Anthropic’s earlier acquisition and storage of millions of pirated digital books

The court found the destructive scanning and AI training use to be fair use under the facts presented. It separately found serious copyright issues connected to the pirated digital collections. In other words, buying a physical book, scanning it, and discarding the original was treated differently from downloading unauthorized digital copies from shadow libraries.

The United States Copyright Office has also emphasized that there is no universal answer for every AI training system. Its Copyright and Artificial Intelligence initiative explains that fair-use questions depend on the details of the use, including the nature of the works, the purpose of copying, the amount used, and the effect on existing or potential markets.

This is a developing area of law, not a blanket permission slip. Future cases may address licensing markets, memorization, market harm, transparency, and whether a model’s outputs compete with the works used to train it. They may also address copyright laws governing derived works, copyrighted content, and other copyrighted works.

Why physical books are more than containers of text

A scanned book can preserve words while losing information that is difficult to reproduce digitally.

A physical book may contain:

  • An edition’s publication date and printing history
  • Marginal notes, stamps, and ownership marks
  • Paper, binding, typography, and cover design
  • Evidence of how a community used or valued the book
  • Physical clues about censorship, repairs, damage, or circulation
  • A material connection to a specific time and place
  • Signatures, pages, and other evidence of ownership or use

A clean digital text may preserve the main content while removing much of this context. If the last physical copy is destroyed, future researchers may know what the words said but not how the object existed in the world.

This is why the debate connects to a broader question of digital ownership. In my guide to open-source philosophy and digital ownership, I explain the difference between using a digital platform and controlling a digital space. The same principle applies here: access to information is not identical to control over its source, history, or future availability.

Is a digital copy enough to preserve knowledge?

Sometimes a digital copy is an excellent preservation tool. Digitization can protect content from fire, flood, decay, dry seasons, and physical loss. It can make rare material searchable and accessible to people who could never visit a particular archive.

But digitization is not automatically preservation. A digital file can be altered, corrupted, deleted, restricted, or trapped inside a private system. Its continued existence depends on storage, backups, file formats, authentication, funding, and the willingness of an organization to maintain access.

A private AI training library creates an additional problem. The public may not be able to inspect the collection, verify what it contains, correct errors, or access the scanned books. If a physical book disappears and the only surviving copy is controlled by a private company, the knowledge may technically exist while becoming practically unavailable.

This is also why I encourage readers to think carefully about control and confidentiality when using AI research tools. My guide to NotebookLM privacy and uploaded documents explains why private does not always mean confidential, especially when information is shared with a hosted AI service.

Photographer: Jonas Jacobsson | Source: Unsplash

What does this mean for authors and publishers?

The book-scanning controversy highlights a difficult imbalance. Authors and publishers create the work, but AI companies may gain enormous commercial value from using collections of that work to build competing products. The original publisher may receive no compensation even when the resulting systems generate higher profit margins for the technology company.

Several questions remain important:

  • Should AI companies license books directly from authors and publishers?
  • Should creators be told whether their work appears in a training dataset?
  • Should there be compensation for commercial training use?
  • Should compensation include renegotiated royalties when AI products become highly profitable?
  • Should libraries and archives receive copies of digitized material?
  • Should rare or nearly unique books receive special protection?
  • Should companies be required to preserve the original artifact when non-destructive scanning is practical?
  • Should a decent hardback copy be preserved when it may be the best surviving edition?

There is no simple answer that protects innovation, creator rights, public access, and historical preservation equally in every case. However, secrecy makes the trade-offs harder to evaluate. A healthier system would provide more transparency about what was acquired, how it was scanned, who controls the resulting files, and whether creators have meaningful choices.

What can individuals and small businesses learn from this?

The story is not only about books. It is also a warning about concentration of control.

When an AI company owns the training library, the model, the infrastructure, and the access rules, users may have little visibility into the origins of the system they rely on. That is similar to the vendor lock-in problem I discuss in practical AI coverage across the Greg Doig technology blog.

A sensible response is not to reject every AI tool. Instead, build more resilient habits:

  • Keep local copies of important documents and research.
  • Use open file formats when possible.
  • Record the sources behind important decisions.
  • Avoid putting sensitive information into AI systems without checking their data policies.
  • Support libraries, archives, authors, and publishers that preserve public knowledge.
  • Use more than one platform for important work.
  • Choose tools that offer export, portability, and clear ownership terms.
  • Ask whether an AI system is genuinely open or merely marketed as open.
  • Be cautious of claims that tools are “copyright-free cars” for information; copyright and ownership still matter.

For people comparing AI assistants, the question should not be only which tool writes the best answer. It should also be where the information goes, who controls the underlying system, and what happens if the provider changes its policies. My guide to ChatGPT alternatives for writing and research explores that question from a practical user perspective.

The real lesson: preservation needs public accountability

AI can help preserve knowledge, but preservation should not depend entirely on private companies whose main goal is to build valuable commercial models.

A responsible preservation system would combine several approaches:

  • Non-destructive scanning for rare and fragile books
  • Public or nonprofit archives with long-term maintenance plans
  • Multiple copies stored in different locations
  • Clear metadata and provenance records
  • Legal access for researchers and the public where appropriate
  • Compensation or licensing systems for creators
  • Independent audits of private digitization projects
  • Preservation of both the text and the physical artifact when possible
  • Good paper, strong signatures, and perfect binding for new archival editions

The central issue is not whether digital copies have value. They do. The issue is whether society should accept irreversible destruction as the default price of building better AI models.

A book can become data without becoming disposable. If companies want access to humanity’s accumulated knowledge, they should be expected to preserve more than the words that can be extracted from it. A durable hardback copy made with bestseller-grade materials should not be treated as the shittiest way to store information simply because a digital file is easier to duplicate.

Photographer: Buddha Elemental 3D | Source: Unsplash

Frequently asked questions

Why are AI companies destroying physical books?

They are using destructive scanning because cutting off the binding allows pages to be scanned quickly and converted into searchable digital files. Those files can then be used for research, data analysis, or AI training.

Did AI companies destroy rare books?

Public reporting and court filings confirm that millions of print books were acquired and destructively scanned. They do not prove that every book was rare, unique, or the last surviving copy. Claims about specific titles should be checked against documented evidence, especially where a book may be the only type available.

Is training AI on books legal?

It depends on the facts and the jurisdiction. In June 2025, a federal court ruled that Anthropic’s use of lawfully purchased books for destructive digitization and AI training was fair use under the circumstances of that case. The same order treated pirated digital books differently.

Is digitization the same as preservation?

No. Digitization preserves access to content, but it may not preserve the physical object, its material history, or public access to the resulting files. Strong preservation usually requires multiple copies, reliable maintenance, clear ownership, and long-term access plans.

How can I protect my own digital knowledge?

Keep local backups, use open formats, document important sources, avoid unnecessary platform dependence, and review how AI services handle uploaded information. Owning a website or domain can also give you more control over your work and audience.

Final thoughts

The most uncomfortable part of this story is not that AI can read books. It is that the race for better training data can turn irreplaceable physical knowledge into a disposable input.

AI may preserve the text while shredding the evidence of how that knowledge lived in the world. Whether that is acceptable depends on what we believe preservation means. If preservation means only that words can be reproduced by a machine, the current process may seem efficient. If preservation includes public access, historical context, creator rights, and the survival of original artifacts, the standard should be much higher.

The question is not simply whether AI is preserving knowledge. The question is who owns the preserved copy, who can verify it, who can access it, and what was destroyed to create it.

For more practical analysis of AI, privacy, digital ownership, and emerging technology, visit GregDoig.com and browse the latest technology articles.