Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

Encyclopedia Britannica sues OpenAI over AI training data

Encyclopedia Britannica and its Merriam‑Webster subsidiary sued OpenAI in Manhattan federal court, alleging the company used their online reference content to train ChatGPT and siphoned web traffic with AI summaries.

The complaint says OpenAI copied nearly 100,000 articles and that ChatGPT can produce near‑verbatim encyclopedia entries and dictionary definitions that divert users.

Britannica also accuses OpenAI of trademark infringement for implying it had permission to reproduce material and for wrongfully citing Britannica in false AI “hallucinations.”

Britannica seeks unspecified monetary damages and a court order to stop the alleged copying.

An OpenAI spokesperson said the company’s models are trained on publicly available data and are grounded in fair use.

The case joins other high‑stakes copyright suits from authors and news outlets and follows Britannica’s ongoing lawsuit against Perplexity AI.

🔗 Source: Reuters

🧠 Food for thought

Implications, context, and why it matters.

The lawsuit comes after licensing talks fell apart

  • Britannica filed the case in the Southern District of New York after licensing negotiations broke down 1.
  • The dispute pits Britannica, a publisher founded in 1768, against OpenAI, a Microsoft-backed AI company with over 900 million weekly users and annual revenue of up to $25 billion 1.
  • The complaint goes beyond copyright. It adds a “false designation of origin” claim tied to AI-generated “hallucinations” (made-up or incorrect outputs) that connect errors to Britannica’s brand and cause reputational harm 1.
  • Court records flag the matter as possibly related to broader multi-district litigation. Multi-district litigation is a process that consolidates similar federal lawsuits to coordinate pretrial proceedings, and it could link this case to wider AI disputes 2.

The case targets how AI answer engines make money

  • The lawsuit, like Britannica’s related action against Perplexity AI, challenges training use plus ChatGPT outputs, including alleged near-verbatim reproductions 3.
  • Britannica argues AI summaries replace publisher products, which cuts traffic and revenue. Some courts have treated direct market substitution as a factor against fair use 45.
  • A win for Britannica could set a precedent for when copyrighted material used to “ground” large language model (LLM) responses through retrieval-augmented generation (RAG) needs a license. RAG is a method that pulls in outside sources to answer a query 5.
  • The result could raise costs for AI answer engines and push more licensing deals. Perplexity AI has signed agreements with some publishers, and OpenAI has done so with other publishers 41.

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.