👩🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔♂️ A friendly human may check it before it goes live. More news here
🧔♂️ A friendly human may check it before it goes live. More news here
Meta sued by publishers over Llama AI training
On May 5, book and academic publishers Elsevier, Cengage, Hachette, Macmillan, and McGraw Hill sued Meta in Manhattan federal court with author Scott Turow, alleging the company used millions of books and journal articles without permission to train its Llama AI models.
The proposed class action says the material included textbooks, scientific papers, and novels such as N.K. Jemisin‘s The Fifth Season and Peter Brown‘s The Wild Robot.
It seeks damages and wider representation for copyright owners.
🔗 Source: Reuters
🧠 Food for thought
Implications, context, and why it matters.
Claims about Meta’s book data add context to the new lawsuit
- Meta has faced earlier claims over Llama training data. In another case, court filings say Meta used LibGen for Llama-related training with approval from CEO Mark Zuckerberg 1. LibGen is a large online collection of free books and papers that critics call pirated 1.
- Plaintiffs also say Meta torrented at least 81.7 terabytes from shadow libraries through Anna’s Archive, a search site that indexes those repositories, including at least 35.7 terabytes from Z-Library and LibGen 2.
- Those moves drew internal objections. One researcher wrote, “I don’t think we should use pirated material,” while engineers worried about torrenting on corporate laptops 2, 3.
- Meta weighed licensing as well, yet one engineering director wrote that licensing even one book would weaken a fair use strategy 2.
Copyright fights around AI now zero in on piracy and data sourcing
- The Anthropic settlement sharpened a legal split between training on lawfully obtained books and getting those books through piracy 4.
- Judge William Alsup ruled that training on legally acquired books counted as fair use, but he rejected summary judgment on piracy claims and found piracy did not qualify as fair use 4.
- Anthropic agreed to a US$1.5 billion settlement after class certification on piracy claims tied to books obtained from LibGen and PiLiMi, another online repository of digitized books 4.
- That approach gives copyright holders a more defined route for claims tied to obtaining and storing pirated copies, beyond the broader dispute over AI training 5.
Recent Meta developments
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.




