Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

Baidu launches Ernie-4.5 AI model to rival OpenAI

Baidu has released a new AI model, ERNIE-4.5-VL-28B-A3B-Thinking, which it claims matches or exceeds the performance of larger models from Google and OpenAI on vision-language benchmarks, despite using fewer computing resources.

Baidu, China’s largest search engine company, said the model can process images, videos, and documents while activating only 3 billion of its 28 billion total parameters through a mixture-of-experts architecture.

The company said that the model’s efficiency allows it to run on a single 80GB GPU, making it accessible for enterprise users.

Baidu claims the model excels in tasks such as document understanding, chart analysis, and visual reasoning, but independent verification of these claims is pending.

The model’s dynamic image analysis feature enables it to zoom in and out for detailed inspection, a function Baidu says is useful for industrial and quality control applications.

ERNIE-4.5-VL-28B-A3B-Thinking is available under the Apache 2.0 license, allowing for commercial use.

🔗 Source: VentureBeat

🧠 Food for thought

Implications, context, and why it matters.

Multimodal reinforcement learning and tooling drive Baidu’s reported gains

  • The model uses GSPO, a reinforcement learning objective for multimodal tasks, and IcePop for data selection 1. Training stays stable across a Mixture-of-Experts setup that routes each input to a small set of experts to cut compute 1. Dynamic difficulty sampling then moves from easier to harder examples 1.
  • Training also targets verifiable tasks using reinforcement learning 1. That departs from typical vision-language training. These are problems with answers that can be auto-checked, such as circuit questions or chart analysis 1.
  • Thinking with Images adds tool use at inference 1. The system can zoom images or run image search to pull info instead of relying only on pre-trained memory 1.
  • Baidu says the lightweight model closely matches results from top flagship models on many benchmarks 1. The docs lack third-party checks, and they offer no direct comparisons to GPT-4V (OpenAI) or Gemini Pro Vision (Google).

Apache 2.0 license and single-GPU option support on-prem visual reasoning services

  • Single-card runs need at least 80GB of GPU memory 1. The release ships under Apache 2.0 for commercial use 1.
  • The repo includes ERNIEKit for supervised fine-tuning, LoRA, and DPO 2. FastDeploy handles inference with OpenAI API compatibility to ease pilot builds 2.
  • Regulated teams can run on-prem for invoice processing or defect detection. This setup fits cases where data cannot leave internal networks.
  • Docs cover usage, though full vLLM support is still in progress, with a “stay tuned” note 3.

Recent Baidu developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.