Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

Wikipedia partners with Microsoft, Meta on AI content training

Wikipedia has unveiled partnerships with Microsoft, Meta, Amazon, Perplexity, and France-based Mistral AI to provide access to its content for AI model training, with some agreements being new and others extensions of previous partnerships.

The Wikimedia Foundation, which operates Wikipedia, said these partnerships are part of its efforts to monetize the use of its data by large tech companies.

The foundation already has a similar arrangement with Google, announced in 2022.

Wikipedia’s 65 million articles in over 300 languages are widely used as training material for generative AI tools.

The nonprofit said that increased use of its content by AI firms has raised server and operational costs, which are mainly covered by public donations.

🔗 Source: Reuters

🧠 Food for thought

Implications, context, and why it matters.

Wikipedia’s AI partnerships bring modest revenue without a big monetization shift

  • Wikimedia Enterprise is the Wikimedia Foundation’s paid data-access service and made $8.3 million in fiscal year 2024-25, which equals 4% of the Foundation’s $185.38 million total 12.
  • Revenue rose 148% year over year, yet the service stays a side income because Foundation policy caps its share at 30% to keep funding mainly from individual donations 1.
  • It became profitable in its fourth year after launch and has netted $646,000 in cumulative profit since January 2022 1.
  • Partners are signing extensions or new deals under the framework launched in 2021, so Wikipedia’s monetization strategy stays the same 3.

AI companies must meet Wikipedia’s attribution rules, which boosts demand for compliance tools

  • The service adds licensing metadata to every API request, which states that over 99.9% of content is under Creative Commons Attribution-ShareAlike licenses that require attribution when reused 4.
  • The APIs add credibility signals like reference_need_score plus reference_risk_score, parsed citations, and structured references that support verifiable sourcing while flagging when an article may need more citations or carries higher reference risk 1.
  • Teams building large language model (LLM) apps, Retrieval-Augmented Generation (RAG) systems (tools that fetch answers from external sources at query time then ground them), or AI knowledge products face growing pressure to provide attribution plus provenance for training data 5.
  • Every article includes complete edit history plus citation trails. That lets provenance-tracking tools trace AI outputs to specific Wikipedia sources with links to community-verified references 6.

Recent Wikipedia developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.