Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

Wikipedia seeks AI licensing deal to cover rising costs

Wikipedia is seeking more licensing agreements with major technology firms to offset rising costs caused by AI companies using its content, co-founder Jimmy Wales said.

The nonprofit Wikimedia Foundation, which operates Wikipedia, signed a deal with Google in 2022 to allow paid access for training AI models and is in talks with other companies.

Wales said automated scraping by AI bots has increased Wikipedia’s server and memory costs, which are not covered by the small public donations the foundation relies on.

While Wikipedia’s content remains free for individual use, for-profit firms training AI models create a disproportionate financial burden.

Wales noted Wikipedia may consider technical tools to limit AI crawling, though this is challenging due to its commitment to open access, and he emphasized the potential impact of public pressure.

🔗 Source: Reuters

🧠 Food for thought

Implications, context, and why it matters.

Wikimedia Enterprise revenue is modest relative to overall funding; AI-related cost pressures are not quantified

  • Wikipedia signed a licensing deal with Google in 2022 and is pursuing others. The Wikimedia Foundation’s 2023-2024 budget says Wikimedia Enterprise, the paid Application Programming Interface (API) and services initiative, generated $3.2 million in annual recurring revenue as of January 2023 1. That equals 1.7% of the $185.3 million in support and revenue the Foundation recorded in 2023 2.
  • Technology makes up 45% of the operating budget for hosting and infrastructure 3. The Foundation ended FY2024 with $271.5 million in net assets 2, yet filings do not break out server-cost jump tied to AI scraping 34. The impact of AI crawling on sustainability remains unclear given flat budget growth 1.

AI developers face unsettled obligations when using CC BY-SA licensed content for training

  • Creative Commons Attribution-ShareAlike (CC BY-SA) licenses require attribution, and they compel sharing adaptations under the same license 5. Some regions allow AI training under copyright exceptions 6. Companies that monetize outputs built on Creative Commons licensed content face murky rules on attribution and whether ShareAlike, the requirement to license derivatives on the same terms, applies 57.
  • This opens a lane for attribution-tracking systems and provenance tooling 5. These tools record data sources and usage history to document compliance during training and inference, when a model generates outputs. Investors can track companies building automated licensing layers that guide developers through ShareAlike duties when model outputs include CC-licensed content 7.

Recent Wikipedia developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.