Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

Authors sue Salesforce for allegedly using books in AI training

Salesforce is facing a proposed class action lawsuit from two authors who say the company used their books without consent to train its AI models.

Molly Tanzer and Jennifer Gilmore allege in a complaint filed on October 13, that Salesforce’s xGen AI models were trained on thousands of copyrighted works, including their own.

The lawsuit claims Salesforce infringed on their rights by using pirated books as training data.

Salesforce declined to comment on the lawsuit.

The authors are represented by attorney Joseph Saveri, who has filed similar lawsuits against other tech firms.

The lawsuit also cites past comments from Salesforce CEO Marc Benioff, who has criticized the use of unlicensed data in AI development.

🔗 Source: Reuters

🧠 Food for thought

Implications, context, and why it matters.

Training data transparency remains elusive for enterprise AI models

  • A lawsuit against Salesforce exposes a gap. The xGen documentation covers architecture and benchmarks 123, yet public materials do not say whether Books3 or similar scraped books fed training.
  • Salesforce trained xGen-7B on 1.5 trillion tokens (small text fragments) from natural language data and code sources 1. It has not published detailed provenance (documented origin) beyond broad categories, so no one can verify claims about pirated content.
  • This opacity clashes with leadership’s past comments on copyrighted training data, which stress appropriate sourcing and payment.
  • Industry reports covered the $1.5 billion Anthropic settlement 4. It puts a price on copyright fights. That raises the risk for firms that have not audited or disclosed their data sources.

Enterprise software vendors face mounting pressure to offer provenance-verified AI

  • AI infrastructure firms can step in with rights-cleared corpora. They can also ship provenance tools. This case surfaces that need for enterprises that want to cut legal risk.
  • Companies building business-to-business (B2B) AI products can offer dataset lineage tracking (records of data origins and transformations) and copyright indemnification (contractual protection against copyright claims) as features. Buyers that deploy models like xGen-7B need assurance they will not inherit vendor liability 5.
  • Recent copyright deals are creating demand for audits that can verify training data provenance after the fact 23. These services help teams gauge exposure before a complaint lands, with a focus on enterprise models.

Recent Salesforce developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.