Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

Ant Group unveils AI framework that is 10x faster than Nvidia’s

Ant Group has released an open-source inference framework called dInfer for diffusion language models, a newer type of AI model that generates outputs in parallel.

The company, based in China and affiliated with Alibaba Group, said dInfer targets models that differ from widely used autoregressive systems like ChatGPT, which generate text one word at a time.

Diffusion models are already popular for image and video generation, and researchers are actively exploring their use in language processing.

Ant claimed dInfer performed up to 3x faster than vLLM, developed by University of California, Berkeley researchers, and 10x faster than Nvidia’s Fast-dLLM, based on internal tests.

The company said dInfer processed 1,011 tokens per second on the HumanEval code-generation benchmark using Ant’s diffusion model, compared to 91 tokens per second for Nvidia’s Fast-dLLM, and 294 for Alibaba’s Qwen-2.5-3B model with vLLM.

🔗 Source: South China Morning Post

🧠 Food for thought

Implications, context, and why it matters.

Diffusion speed claims lack independent quality checks

  • Ant’s dInfer claims 10x over Nvidia’s Fast-dLLM and 3x over vLLM. On HumanEval (a widely used code benchmark), Ant measured 1,011 tokens per second on LLaDA-MoE (a diffusion language model using a Mixture of Experts), versus 91 for Fast-dLLM and 294 for Alibaba’s Qwen-2.5-3B (a 3-billion-parameter language model served with the open-source vLLM engine).
  • Perplexity (how well a model predicts the next token), Massive Multitask Language Understanding (MMLU) accuracy, and MT-Bench (a multi-turn dialogue evaluation) would test whether diffusion models match autoregressive (AR) LLMs on task performance. They are absent today.
  • LLaDA says it can rival LLaMA 3 8B (an 8-billion-parameter open model from Meta) when trained from scratch 1. Apple researchers say diffusion language models (DLMs) remain small in scale, and lack fair comparisons on language modeling benchmarks 2.
  • Without third-party checks on quality and speed trade-offs, claims of superiority over AR models remain preliminary, not proven.

Inference providers can build diffusion serving layers

  • Diffusion language models enable parallel generation stacks, unlike sequential decoding in autoregressive (AR) inference.
  • API and hosting vendors could ship diffusion endpoints. Targets include code completion or draft generation where speed matters more than reasoning depth.
  • Open-source options include Dream 7B (a 7-billion-parameter diffusion language model), Open-dCoder (a diffusion-based code model), and LLaDA (a family of diffusion LLMs) 341. Repositories tracking 13 public diffusion LLM projects offer implementation references 5.
  • Upside depends on adoption beyond research. Providers should watch benchmark results and enterprise demand before big bets.

Recent Ant Group developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.