Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

Nvidia’s Blackwell chips double AI training speed: report

Nvidia’s latest Blackwell chips have shown notable improvements in training large AI systems, according to benchmark data released by MLCommons on June 4, 2025.

MLCommons, a nonprofit organization focused on evaluating AI performance, reported that the Blackwell chips are more than twice as fast as the previous-generation Hopper chips on a per-chip basis.

The benchmark tests included training AI models like Meta’s Llama 3.1 405B, which involves complex tasks with trillions of parameters.

In one test, 2,496 Blackwell chips completed the training in 27 minutes.

The older Hopper chips required over three times as many units to achieve a faster time.

🔗 Source: Reuters


🧠 Food for thought

1️⃣ The efficiency revolution in AI chip design accelerates

Nvidia’s Blackwell chips represent a significant leap in training efficiency, reflecting a broader industry trend toward doing more with less computational resources.

The data shows Blackwell chips are more than twice as fast per chip as the previous Hopper generation, enabling 2,496 Blackwell chips to complete training tasks that previously required three times as many chips 1.

This efficiency focus aligns with developments across the semiconductor industry, where companies have been designing specialized AI chips that optimize performance for specific workloads rather than general computing tasks 2.

The trend toward efficiency isn’t exclusive to Nvidia—companies like Google with its Tensor Processing Units (TPUs) and Amazon with Trainium chips have been pursuing similar goals, recognizing that AI’s computational demands require specialized architectures 1.

These advancements are particularly crucial as AI models continue to grow in size and complexity, with the industry seeking sustainable approaches to the enormous computational requirements of advanced AI systems.

2️⃣ AI infrastructure evolving from monolithic systems to specialized clusters

CoreWeave’s observation about companies using smaller groups of chips for separate AI training tasks represents a fundamental architectural shift in how AI infrastructure is deployed.

Rather than building homogenous systems with 100,000+ identical chips, organizations are increasingly developing specialized subsystems optimized for specific portions of the AI training process 3.

Recent Nvidia developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.