- Premium Content It takes our newsroom weeks - if not months - to investigate and produce stories for our premium content. You can’t find them anywhere else.
Frontier dreams stall as labs turn to hosting Chinese models
Sarvam is India’s only hope to join the ranks of the AI giants. But at its first flagship event last month, the company did not unveil the model meant to take it there.
Instead, Sarvam announced that it would use its GPUs to help Indian companies run open models made by some of China’s most formidable AI labs.

Image credit: Ulla
“The Chinese models have become very competitive,” said Sarvam co-founder Vivek Raghavan at a press conference during the company’s Epoch event last month. “To be able to serve the best models in the world at the most competitive prices … in India is very critical.”
At the same event, Sarvam announced that its latest flagship model was still six months away. While the company stopped short of saying so, it is becoming clear that Sarvam is part of a growing group of AI labs that are re-selling their compute capacity by hosting open source models.
“We will ensure that we can aggregate GPUs in the country,” says Pratyush Kumar, CEO of Sarvam. “We also have maybe the largest set of [Nvidia] Blackwells in the country. We have quite a bit of compute, which is required to be able to serve this.” The company has 2,000 Nvidia Blackwell GPUs in India and plans to expand that to 10,000.
France’s Mistral is following a similar path. Mistral Compute sells European GPU capacity and enterprise AI services, and it now hosts models from China’s Z.ai alongside its own.
US-based Zyphra has also launched an inference platform for open-weight models from DeepSeek, Kimi, and other players, running on AMD GPUs supplied by neocloud TensorWave.
Both Sarvam and Mistral have been attempting to bring out proprietary closed models as a pitch, only to find Chinese labs producing near-frontier open models that customers can download for free and host themselves.
“Why would an enterprise want to buy proprietary sovereign models without the frontier capabilities when you can deploy near frontier tech at a much lower cost?” says Manoj Sukumaran, an independent analyst with expertise in data centers and compute. He sees these labs’ move to host rival open models as an attempt to create an early revenue stream.
The scarcity paradox
These new inference businesses are a consequence of the way AI compute is bought.
High-end GPUs usually go to data centers, particularly neoclouds. These firms usually rent them out under long-term contracts with clients, according to Mahesh Kolli, co-founder and president of neocloud firm AM Intelligence.
Faced with a global compute shortage and a backlog of Nvidia orders, AI labs rent these chips in anticipation of expected training needs. But with the rising cost of training as well as growing competition from Chinese models, many of these chips are sitting idle.
According to a report by CastAI, hyperscaler facilities only use 5% of their GPUs’ capacity.
Real estate of GPUs
Will they want the compute back?
New chips, new economics
Stay ahead in Asia’s tech landscape
This is premium content. Subscribe to read the full story.
Are AI firms pivoting to cloud providers? Not quite, but many are finding ways to monetize their idle GPUs
We know this is not ideal. ⌛ Sign up in 20 seconds. Cancel anytime.
Our subscriber community includes professionals from these companies:





Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.

