Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

DeepSeek unveils AI self-verifying math reasoning model

DeepSeek, a Chinese AI firm, has released DeepSeekMath-V2, an open-source mathematical reasoning model.

The model is available on Hugging Face and GitHub and uses a self-verifying framework where two large language models work together, with one generating proofs and the other reviewing them.

DeepSeekMath-V2 achieved scores on the 2025 International Mathematical Olympiad and 2024 Chinese Mathematical Olympiad that matched top results, and it scored 118 out of 120 on the 2024 Putnam Exam, above the highest human score of 90.

The model outperformed DeepMind’s DeepThink in the IMO-ProofBench benchmark.

DeepSeek said the self-verifying approach is intended to ensure both correct answers and sound reasoning.

The company claims this addresses a key challenge in AI-driven mathematical reasoning, where correct answers do not always mean correct logic.

🔗 Source: Xinhua

🧠 Food for thought

Implications, context, and why it matters.

DeepSeekMath-V2’s scores require external verification to separate breakthrough from benchmark overfitting

  • DeepSeekMath-V2 scored 118/120 on Putnam and posted gold-medal IMO results. Commenters say 2024 Putnam problems appeared in some models’ Reinforcement Learning (RL) training data 1, which creates contamination risk (evaluation problems in training corpora).
  • The earlier DeepSeekMath reached 51.7% on the MATH benchmark (a standardized dataset of problems for evaluating language models) without external toolkits or voting techniques 2 (e.g., plug-in solvers and answer-ensemble or majority-vote prompts). DeepSeekMath-V2’s competition results need independent replication by academic benchmark maintainers or third-party audits to check for genuine reasoning and to rule out leakage.
  • DeepSeekMath-V2 was tested on IMO-ProofBench. One benchmark does not settle performance across varied math domains, and proof skills may not carry over to creative idea generation where large language models still struggle 1.

Cloud providers can offer dual-LLM verification stacks for math-intensive applications

  • A verifier and generator setup (two models where one produces a solution and the other checks it) drives demand for hosted, low-latency dual-LLM endpoints that scale verification compute 3.
  • DeepSeekMath-V2 has 685B parameters and a 689GB footprint that demands heavy GPU capacity 4. Providers can ship optimized inference stacks with custom NVIDIA Compute Unified Device Architecture (CUDA) kernels, echoing DeepSeek-V2 training tweaks 5, plus quantization options (reduced-precision model weights) and VRAM or throughput Service-Level Agreements (SLAs).
  • Apache 2.0 permits commercial use 6, so MLOps startups can ship vertical solutions for finance (quantitative analysis verification) or pharmaceuticals (computational chemistry validation) that need step-by-step, provable reasoning chains beyond final answers.

Recent DeepSeek developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.