Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

Nvidia servers boost China’s Moonshoot AI models by 10x

Nvidia said its new GB200 NVL72 rack-scale AI server can speed up open-source mixture-of-experts (MoE) models by as much as 10x.

The improvement was measured using models such as Moonshoot AI’s Kimi K2 Thinking compared with the older HGX H200 platform.

MoE architectures split AI tasks among specialized components, activating only the most relevant ones for each token, which makes them more efficient than traditional dense models.

Nvidia says its GB200 NVL72 system links 72 Blackwell GPUs through high-speed NVLink connectivity, enabling MoE models to scale across more GPUs and reducing memory and communication bottlenecks.

The company claims this design allows for faster inference, lower compute demands, and supports longer input lengths.

Major cloud providers, including AWS, Google Cloud, Microsoft Azure, and CoreWeave, are deploying the GB200 NVL72 to support enterprise AI workloads.

🔗 Source: Nvidia

🧠 Food for thought

Implications, context, and why it matters.

GB200 NVL72 availability remains unclear despite performance claims

  • Nvidia cites 10x speedups for MoE models like Kimi K2, yet GB200 NVL72 pricing is undisclosed and rollout timing differs by provider, muddying 2025 capacity plans 12.
  • Analysts cut 2025 shipment forecasts for GB200 NVL72 cabinets by over 50% to 25,000 to 35,000 units, equaling 2.52 million GPUs 3.
  • Together AI, an AI infrastructure provider, says delivery takes 4 to 6 weeks with no “NVIDIA lottery” (industry shorthand for constrained GPU allocation queues), with full‑rack NVL72 clusters ready now 1. Nebius, a European AI‑focused cloud platform, requires contracts signed at least a month before the start date 2.
  • Rumors about GB300 and B300 GPUs with 50% more power landing about six months after the 200 series may push buyers to wait 3.

Software vendors can build managed MoE services as open-source models mature

  • For software integrators plus cloud builders, GB200 NVL72 has 130 TB/s NVLink (Nvidia’s GPU interconnect) bandwidth and up to 13.4 TB of High Bandwidth Memory (HBM) 3e 1. That enables open MoE options like Qwen3‑235B‑A22B for production use 4. Qwen3‑235B‑A22B has 235 billion total parameters with roughly 22 billion active per token, and DeepSeek V3.1 is an open MoE 4.
  • These models use permissive licenses such as Apache 2.0 or MIT that allow commercial deployment 54. Their workloads are sparse, so cost tracks the number of active experts rather than total parameters 6.
  • Teams can build cost‑performance tools that use sparsity‑aware metrics such as S‑MBU and S‑MFU to better measure memory or compute use in production 6.
  • Open models now support context windows up to 128K tokens 5. That unlocks translation, coding help, or long‑document reasoning without costly proprietary systems.

Recent NVIDIA developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.