Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

Meta sued by publishers over Llama AI training

On May 5, book and academic publishers Elsevier, Cengage, Hachette, Macmillan, and McGraw Hill sued Meta in Manhattan federal court with author Scott Turow, alleging the company used millions of books and journal articles without permission to train its Llama AI models.

The proposed class action says the material included textbooks, scientific papers, and novels such as N.K. Jemisin‘s The Fifth Season and Peter Brown‘s The Wild Robot.

It seeks damages and wider representation for copyright owners.

The case adds to a broader fight over whether AI training on copyrighted material counts as fair use.

Creators have also sued OpenAI and Anthropic, while courts have issued mixed early rulings and Anthropic settled one case for US$1.5 billion last year.

🔗 Source: Reuters

🧠 Food for thought

Implications, context, and why it matters.

Claims about Meta’s book data add context to the new lawsuit

  • Meta has faced earlier claims over Llama training data. In another case, court filings say Meta used LibGen for Llama-related training with approval from CEO Mark Zuckerberg 1. LibGen is a large online collection of free books and papers that critics call pirated 1.
  • Plaintiffs also say Meta torrented at least 81.7 terabytes from shadow libraries through Anna’s Archive, a search site that indexes those repositories, including at least 35.7 terabytes from Z-Library and LibGen 2.
  • Those moves drew internal objections. One researcher wrote, “I don’t think we should use pirated material,” while engineers worried about torrenting on corporate laptops 2, 3.
  • Meta weighed licensing as well, yet one engineering director wrote that licensing even one book would weaken a fair use strategy 2.

Copyright fights around AI now zero in on piracy and data sourcing

  • The Anthropic settlement sharpened a legal split between training on lawfully obtained books and getting those books through piracy 4.
  • Judge William Alsup ruled that training on legally acquired books counted as fair use, but he rejected summary judgment on piracy claims and found piracy did not qualify as fair use 4.
  • Anthropic agreed to a US$1.5 billion settlement after class certification on piracy claims tied to books obtained from LibGen and PiLiMi, another online repository of digitized books 4.
  • That approach gives copyright holders a more defined route for claims tied to obtaining and storing pirated copies, beyond the broader dispute over AI training 5.

Recent Meta developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.