Recursos

en · Inglés

$1.5 billion for pirated books: the price of training data

Anthropic agrees to pay authors $1.5 billion, around $3,000 per work. Not for training on books, but for how it obtained them. The distinction matters — a great deal.

Por Carolina Buldain · Partner de NOVAZZ Publicado el 16 de septiembre de 2025

On 5 September, news broke of the settlement to close the class action brought by a group of authors against Anthropic, the company that develops Claude: $1.5 billion, plus interest, at around $3,000 per work for roughly half a million books. It has been described as the largest copyright recovery in US history. The settlement is pending the judge’s approval.

The details are on the official settlement website and in the Authors Guild’s guide for authors.

The key distinction

To understand the case, we need to go back to June. Judge William Alsup issued a ruling at the time with two very different parts:

  • Training a model on lawfully purchased books was, in his view, “exceedingly transformative” and therefore protected as fair use.
  • Building a central library with copies downloaded from pirate sites was not.

The settlement resolves that second part. As well as paying, Anthropic must destroy the pirated files. It does not cover what the models may generate in future.

What I find relevant

The case can be read as a partial victory for each side. For the AI industry, the principle that training on lawfully acquired works can be legitimate is important. For authors, the message is just as clear: where the data comes from matters, and skipping that step comes at a price.

From an ethical standpoint, that is what I take away. For years, the idea has spread that everything on the internet is free raw material. This settlement does not settle the debate, but it puts a figure — around $3,000 — on the value of each work used without permission.

And for companies that use AI

Few companies train models, but all of them choose providers. And choosing a provider also means choosing its practices. Some reasonable questions:

  1. What does the provider say about the origin of its training data?
  2. Does it offer contractual guarantees against intellectual property claims?
  3. Does it use our data to train its models? Can we prevent it?

These are not questions for a multinational’s legal department. They are the same ones we would ask any supplier about the origin of what it sells us.

Primer paso

¿Quieres aplicarlo en tu empresa?

Cuéntanos tu caso en una llamada de 30 minutos. Sin compromiso.

  1. 01Nos escribesDos o tres líneas bastan. Sin formularios eternos ni llamadas comerciales.
  2. 02Te respondemos en 24 hCon preguntas concretas sobre tu caso, no con un catálogo.
  3. 03Si encaja, diagnóstico gratuito30 minutos para ver dónde la IA tiene sentido y por dónde empezar.

¿Prefieres tu propio correo? Escríbenos a info@novazz.es

Sin compromiso · Respuesta en 24h · 100% confidencial

Responsable: ZUAZO SOLUTIONS, S.L.. Finalidad: atender tu solicitud y, si lo pides, agendar una sesión de diagnóstico. Derechos: acceso, rectificación, supresión, oposición y otros, escribiendo a info@novazz.es. Más información en la Política de Privacidad.