Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

Reddit sues Perplexity for allegedly scraping its data

Reddit has filed a lawsuit against Perplexity, accusing the AI startup of unlawfully scraping its data to train its AI-powered search engine.

The case, filed in a New York federal court, also names Lithuania-based Oxylabs, Russia-based AWMProxy, and Texas-based SerpApi as defendants.

Reddit claims these companies bypassed its data protection systems and alleges Perplexity worked with at least one of them to access Reddit content without a license.

Reddit said it sent a cease-and-desist letter to Perplexity last year, but citations to Reddit increased 40x afterward.

Perplexity said it “will not tolerate threats against openness and the public interest” and intends to defend itself in court.

SerpApi and Oxylabs also plan to defend themselves, with Oxylabs saying it was “shocked and disappointed” by the lawsuit.

Reddit is seeking unspecified damages and a court order blocking Perplexity from using its data.

🔗 Source: Reuters

🧠 Food for thought

Implications, context, and why it matters.

Reddit’s lawsuit reveals the legal doctrine gap in data scraping cases

  • Reddit leans on Digital Millennium Copyright Act (DMCA) Section 1201, an anti-circumvention rule that bars bypassing access controls 1. It alleges scrapers ignored robots.txt (a site file that tells automated crawlers what they may index), IP rate limits (caps on requests per Internet Protocol address), Completely Automated Public Turing test to tell Computers and Humans Apart (CAPTCHA) systems, and Google’s SearchGuard protections (Google Search anti-bot/abuse defenses) 1.
  • Its case turns on proof of active evasion rather than routine access to public data 1. The complaint logs nearly 3 billion search engine results pages in two weeks of July 2025, including 1.05 billion scraped by SerpApi (a search results API provider) over seven days using “Ludicrous Speed Max” to dodge CAPTCHA and IP blocks 1.
  • Anthropic’s $1.5 billion copyright settlement steers AI firms toward licensing 2.

Compliance monitoring services become critical for AI companies navigating scraping litigation

  • More suits raise demand for third-party compliance checks on training data. Reddit planted a “test post”, reachable only through Google Search, to catch Perplexity pulling content without permission 1.
  • Cloudflare (a web infrastructure and security company) launched AI bot blocking and a Pay Per Crawl marketplace that lets automated agents pay site owners for access 2.
  • AI model teams need vendor diligence on data acquisition, with indemnification and proof of lawful sourcing. SerpApi lists Perplexity as a customer 1, which raises exposure. Licensing for grounding or Retrieval-Augmented Generation (RAG) gives a safer path 3. Perplexity shares 80% of related revenue with publishers 2, and Gannett, a large U.S. newspaper chain, joined its Publisher Program 3.

Recent Reddit developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.