Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

Alibaba said to order over 40,000 AMD AI chips

Alibaba is considering purchasing 40,000 to 50,000 MI308 AI accelerators from AMD, according to sources familiar with the matter.

The MI308, tailored by AMD for the China market and approved for US export, requires a 15% licensing fee to US authorities.

The chip features 192GB of HBM3 memory and can run long-context inference for 70-billion-parameter large language models on a single card.

It is priced at around US$12,000, about 15% cheaper than Nvidia’s H20 chip.

The MI308 has not faced the same security scrutiny as some other chips.

🔗 Source: TechNode

🧠 Food for thought

Implications, context, and why it matters.

MI308’s real advantage may be memory, not price

  • Reports peg the MI308 at 15% less than Nvidia’s H20. Its 192GB of High Bandwidth Memory (HBM3) enables single-card runs of 70-billion-parameter models, with MI308 inference cited 1.
  • That capacity eases long-context inference. Enterprises can skip multi-GPU setups, which cuts costs and engineering work for distributed inference.
  • Both H20, Nvidia’s export-limited China-market GPU, and MI308 face export caps. H20’s 96GB may require model sharding (splitting the model across multiple GPUs) for large-context 70B workloads, while MI308’s headroom enables simpler deployment for document processing or extended dialogue systems 1.
  • AMD’s ROCm software stack, its open-source GPU compute platform, trails Nvidia’s CUDA GPU programming platform. Enterprises may need porting and optimization work, which can erase hardware savings unless Microsoft’s CUDA-to-ROCm toolkits reach production 2.

Infrastructure vendors can capture migration revenue if Alibaba’s order materializes

  • If Alibaba buys 40,000 to 50,000 MI308s, then deploys them, demand rises for help on CUDA-to-ROCm migration plus inference tooling or hybrid GPU orchestration.
  • Microsoft is building CUDA-to-ROCm translation toolkits 2, which hints at need for conversion layers that cut code rewrites. Consultancies and tool vendors can package these for on-premise Chinese deployments outside Azure, Microsoft’s cloud.
  • Cloud infrastructure providers and system integrators are firms that assemble hardware, customize software. They can win deals with stacks that hide AMD and Nvidia differences. Customers hedge future export limits while keeping parity across GPU types.
  • A 40,000 to 50,000 unit plan would push Alibaba to build ROCm deployment skills in-house. That blueprint could spread to smaller Chinese clouds and AI labs with vendor help, growing the market beyond Alibaba.

Recent AMD developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.