👩🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔♂️ A friendly human may check it before it goes live. More news here
🧔♂️ A friendly human may check it before it goes live. More news here
China’s Moonshot builds AI with fewer advanced chips than US rivals
Moonshot AI, a Beijing-based AI startup, says it develops models with fewer high-end GPUs compared to US competitors.
In a Reddit session, a company executive using the handle “ppwwyyxx” said Moonshot is “outnumbered” by US firms in terms of high-end GPUs available for model training.
The executive confirmed that Kimi K2 Thinking, a reasoning-focused version of the company’s Kimi K2 model, was trained using Nvidia’s older H800 GPUs.
The H800 chips were banned from export to China in late 2023.
🔗 Source: South China Morning Post
🧠 Food for thought
Implications, context, and why it matters.
Kimi K2 Thinking tops benchmarks on older Nvidia H800 chips
- Moonshot AI trained Kimi K2 Thinking on Nvidia H800 GPUs banned from export to China in late 2023, yet it beats leading closed models on Humanity’s Last Exam (a general reasoning test) and BrowseComp (a web-browsing agent benchmark) 1.
- 1 trillion parameter model uses 32 billion active parameters (meaning a subset of parameters is used per token), and applies Quantization-Aware Training with INT4 (4-bit integer) weight-only quantization for about 2x faster generation (inference throughput) while maintaining top performance 1.
- These gains suggest training and routing can offset limited GPUs. Kimi K2 Thinking runs 200–300 sequential tool calls (invocations of external tools such as a browser or code interpreter or retrieval system) without human help while staying competitive 1.
Software gaps on Chinese accelerators open room for infrastructure vendors
- US curbs on newer Nvidia chips push Chinese labs to options like Huawei’s Ascend 910B-based Atlas A2/A3 systems (servers built around the Ascend AI accelerator), and frameworks like vllm-ascend (an adaptation of the vLLM inference engine for Ascend hardware) target these platforms 2.
- vllm-ascend lists 910 and 910 Pro B (earlier Ascend variants) as ‘unplanned yet’ for support 2, while users request guidance for running DeepSeek v3 (a large language model by DeepSeek) on 910B machines 3.
- Vendors or systems integrators that port LLM serving stacks (software to deploy models in production), quantization tools, or optimization frameworks to Ascend hardware could meet enterprise demand for alternatives to Nvidia-dependent infrastructure 2.
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.




