Tired of ads? Enjoy an ad-free experience by signing up.
  • Premium Content
    It takes our newsroom weeks - if not months - to investigate and produce stories for our premium content. You can’t find them anywhere else.
Pranav Balakrishnan · · 6 min read

Frontier dreams stall as labs turn to hosting Chinese models

Sarvam is India’s only hope to join the ranks of the AI giants. But at its first flagship event last month, the company did not unveil the model meant to take it there.

Instead, Sarvam announced that it would use its GPUs to help Indian companies run open models made by some of China’s most formidable AI labs.

Image credit: Ulla

“The Chinese models have become very competitive,” said Sarvam co-founder Vivek Raghavan at a press conference during the company’s Epoch event last month. “To be able to serve the best models in the world at the most competitive prices … in India is very critical.”

At the same event, Sarvam announced that its latest flagship model was still six months away. While the company stopped short of saying so, it is becoming clear that Sarvam is part of a growing group of AI labs that are re-selling their compute capacity by hosting open source models.

“We will ensure that we can aggregate GPUs in the country,” says Pratyush Kumar, CEO of Sarvam. “We also have maybe the largest set of [Nvidia] Blackwells in the country. We have quite a bit of compute, which is required to be able to serve this.” The company has 2,000 Nvidia Blackwell GPUs in India and plans to expand that to 10,000.

France’s Mistral is following a similar path. Mistral Compute sells European GPU capacity and enterprise AI services, and it now hosts models from China’s Z.ai alongside its own.

US-based Zyphra has also launched an inference platform for open-weight models from DeepSeek, Kimi, and other players, running on AMD GPUs supplied by neocloud TensorWave.

Both Sarvam and Mistral have been attempting to bring out proprietary closed models as a pitch, only to find Chinese labs producing near-frontier open models that customers can download for free and host themselves.

“Why would an enterprise want to buy proprietary sovereign models without the frontier capabilities when you can deploy near frontier tech at a much lower cost?” says Manoj Sukumaran, an independent analyst with expertise in data centers and compute. He sees these labs’ move to host rival open models as an attempt to create an early revenue stream.

The scarcity paradox

These new inference businesses are a consequence of the way AI compute is bought.

High-end GPUs usually go to data centers, particularly neoclouds. These firms usually rent them out under long-term contracts with clients, according to Mahesh Kolli, co-founder and president of neocloud firm AM Intelligence.

Faced with a global compute shortage and a backlog of Nvidia orders, AI labs rent these chips in anticipation of expected training needs. But with the rising cost of training as well as growing competition from Chinese models, many of these chips are sitting idle.

According to a report by CastAI, hyperscaler facilities only use 5% of their GPUs’ capacity.

Real estate of GPUs

Will they want the compute back?

New chips, new economics

Stay ahead in Asia’s tech landscape

This is premium content. Subscribe to read the full story.

Why subscribe?

Are AI firms pivoting to cloud providers? Not quite, but many are finding ways to monetize their idle GPUs

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

10

10 company database access

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

🧠 For professionals / ⭐ Best value

CoreBest value

US$16.58US$14.92/month

Billed annually at US$179.10 on the first year

Get instant access to this article and more every month

Unlimited premium content

Unlimited news briefs & articles

Unlimited company database access

Ad-free reading experience

Just US$0.55 per day

Save US$19.90 on the first year. Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.

TIA Writer

Pranav Balakrishnan

AI beat reporter based in Bengaluru, India. If you’re building in AI, investing in it, researching it, or simply navigating its impact in your work or life, I’d love to hear from you.