Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

China’s Moonshot builds AI with fewer advanced chips than US rivals

Moonshot AI, a Beijing-based AI startup, says it develops models with fewer high-end GPUs compared to US competitors.

In a Reddit session, a company executive using the handle “ppwwyyxx” said Moonshot is “outnumbered” by US firms in terms of high-end GPUs available for model training.

The executive confirmed that Kimi K2 Thinking, a reasoning-focused version of the company’s Kimi K2 model, was trained using Nvidia’s older H800 GPUs.

The H800 chips were banned from export to China in late 2023.

🔗 Source: South China Morning Post

🧠 Food for thought

Implications, context, and why it matters.

Kimi K2 Thinking tops benchmarks on older Nvidia H800 chips

  • Moonshot AI trained Kimi K2 Thinking on Nvidia H800 GPUs banned from export to China in late 2023, yet it beats leading closed models on Humanity’s Last Exam (a general reasoning test) and BrowseComp (a web-browsing agent benchmark) 1.
  • 1 trillion parameter model uses 32 billion active parameters (meaning a subset of parameters is used per token), and applies Quantization-Aware Training with INT4 (4-bit integer) weight-only quantization for about 2x faster generation (inference throughput) while maintaining top performance 1.
  • These gains suggest training and routing can offset limited GPUs. Kimi K2 Thinking runs 200–300 sequential tool calls (invocations of external tools such as a browser or code interpreter or retrieval system) without human help while staying competitive 1.

Software gaps on Chinese accelerators open room for infrastructure vendors

  • US curbs on newer Nvidia chips push Chinese labs to options like Huawei’s Ascend 910B-based Atlas A2/A3 systems (servers built around the Ascend AI accelerator), and frameworks like vllm-ascend (an adaptation of the vLLM inference engine for Ascend hardware) target these platforms 2.
  • vllm-ascend lists 910 and 910 Pro B (earlier Ascend variants) as ‘unplanned yet’ for support 2, while users request guidance for running DeepSeek v3 (a large language model by DeepSeek) on 910B machines 3.
  • Vendors or systems integrators that port LLM serving stacks (software to deploy models in production), quantization tools, or optimization frameworks to Ascend hardware could meet enterprise demand for alternatives to Nvidia-dependent infrastructure 2.

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.