Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

Wikipedia urges AI firms to use paid platform over scraping

Wikipedia is urging AI companies to use its paid Wikimedia Enterprise platform instead of scraping its website for content.

The Wikimedia Foundation said the platform helps prevent server overload and supports its nonprofit mission.

The foundation reported a surge in automated traffic in May and June as some AI bots tried to bypass detection systems, while human page views dropped 8% year-on-year.

It also asked AI firms to properly attribute its content and highlight original sources.

Earlier this year, the foundation outlined an AI strategy focused on tools to support editors rather than replace them.

🔗 Source: TechCrunch

🧠 Food for thought

Implications, context, and why it matters.

Wikimedia Enterprise pricing and voluntary uptake amid AI scraping

  • Enterprise offers a free tier until data egress (data transferred out of Wikimedia systems) exceeds limits 1. Google and the Internet Archive are customers 2. The latter gets it at no cost as a nonprofit, while Google’s terms remain undisclosed 2. Use is opt-in, and companies can keep using free tools 3.
  • The project covered operating costs within a year of the 2021 launch 2. Paid customers get a 99% Service Level Agreement (SLA) with stronger interfaces, while AI firms can still use public Application Programming Interfaces (APIs) with data dumps under Wikimedia’s open-licensing model 3. This approach differs from Reddit’s reported $60 million yearly licensing deal with Google 4.

Bot management vendors see new demand as AI scraping spreads across publishers

  • Bytespider and Amazonbot lead AI crawler traffic 4. ClaudeBot and GPTBot sit close behind 4. AI bots hit 39% of top sites, yet only 3% block them, which creates room for services from Cloudflare and Fastly 4.
  • Some AI operators spoof user agents (the identifier strings browsers send to websites) 4. Teams use machine learning-based fingerprinting (analyzing multiple technical signals to identify a client) beyond robots.txt rules (a standard file that tells bots what they may crawl) 4. Fastly’s AI Bot Management and Cloudflare’s single-click blocking feature target it 54. SaaS companies and content platforms can add tiered controls with real-time monitoring, then monetize training access through pay-per-crawl arrangements (structured agreements in which crawlers pay per request or dataset access) 6.

Recent Wikimedia developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.