Viral Wire

AI Training on Copyrighted Books: Google, Anthropic, OpenAI in Legal Crosshairs

Google's 25M scanned books sit unused; Anthropic settles; court says no permission needed.

Deep Dive

The AI industry's appetite for training data is colliding with copyright law, as multiple high-profile cases expose how tech giants treat books. Google, despite digitizing approximately 25 million books (many out-of-print) through its library project, keeps them locked away with no public access — a vast, idle dataset that could theoretically train large language models. Meanwhile, Anthropic settled a class-action lawsuit with authors who claimed the company used pirated books to train its Claude model; however, some writers have refused the settlement, signaling deeper discontent. Compounding the complexity, a recent US court ruling stated that AI companies can train on legally obtained books without seeking individual authors' permission, a decision that could reshape the industry's data practices.

OpenAI, too, is under scrutiny for allegedly removing a dataset of pirated books from its training corpus without explaining why. These episodes underscore the regulatory vacuum around AI training data. While tech giants publicly endorse open AI models and responsible data use, their actions reveal a reliance on copyrighted content. The outcome of these legal battles — combined with legislative pushes in the US and EU — will determine whether AI developers can freely use existing creative works or must negotiate licensing agreements. For authors and publishers, the stakes are existential: their livelihoods hinge on control over their intellectual property in an era of automated content generation.

Key Points
  • Google holds 25M scanned books (including rare titles) but provides no public access or AI training use.
  • Anthropic settled with authors over pirated-book training for Claude; some authors reject the deal, seeking court-appointed oversight.
  • A US court ruled that legally acquired books do not need author permission for AI training, potentially setting a precedent.

Why It Matters

These cases will shape copyright law for AI, affecting data access for training and author compensation globally.

📬 Get the top 10 AI stories daily