Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

OpenAI launches IndQA to test AI on Indian culture, languages

OpenAI has launched IndQA, a new benchmark to evaluate AI models on questions rooted in Indian languages and culture.

IndQA includes 2,278 questions in 12 languages across 10 cultural domains, developed with input from 261 Indian experts.

The benchmark tests AI systems on reasoning and context in areas such as literature, food, history, law, and sports.

Unlike previous benchmarks, it focuses on culturally relevant, complex questions rather than just translation or multiple-choice tasks.

OpenAI said the questions were designed to be difficult for leading AI models, including GPT-4o and GPT-5, with only those that models failed to answer retained.

Each question includes a grading rubric and an ideal response based on expert expectations.

Performance scores show even top models score below 40%, highlighting ongoing challenges in non-English language understanding.

IndQA is intended to track AI improvement over time and help develop benchmarks for other low-resource languages.

🔗 Source: OpenAI

🧠 Food for thought

Implications, context, and why it matters.

IndQA’s public accessibility determines if it becomes an industry standard or just internal marketing

  • This depends on public release of IndQA’s dataset, rubrics, and evaluation code with clear licensing. A leaderboard that accepts third-party submissions and posts a transparent ranking matters. SimpleQA Verified offers datasets, evaluation code, and leaderboards on Kaggle (a public data science platform) 1. The Open-LLM-Leaderboard posts results on HuggingFace (an open hub for AI models and evaluations) 2.
  • Without public access and independent checks, IndQA will get tagged as PR, not a trusted yardstick. The Foundation Model Leaderboards repository (a curated list that sets inclusion criteria for trustworthy leaderboards) sets inclusion rules 3, and the community expects transparent, reproducible benchmarking.

India-based data vendors face immediate demand as top models score below 40% on IndQA

  • Teams building for Indian markets need Indian-language data collection and annotation (human labeling of examples). They also need Reinforcement Learning from Human Feedback (RLHF) and evaluation services to lift IndQA scores. This opens partnership and investment paths with India-based vendors that supply culturally grounded training data across 12 languages plus 10 cultural domains in IndQA.
  • AI platform and data operations groups should find regional labeling firms that grasp local nuance. Investors can target roll-ups (consolidating smaller firms through acquisition) of fragmented language service providers. The 261 Indian experts who contributed to IndQA mirror the talent pool needed to meet demand for culturally relevant AI training data.

Recent OpenAI developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.