🧔♂️ A friendly human may check it before it goes live. More news here
Alibaba’s AI tops global math contests, challenging OpenAI
Alibaba has unveiled Qwen3-Max-Thinking, an upgraded AI model that achieved perfect scores in the 2025 American Invitational Mathematics Examination (AIME) and the Harvard-MIT Mathematics Tournament (HMMT).
Alibaba said this is the first time a Chinese AI reasoning model has reached 100% accuracy in these contests.
These competitions are considered among the most difficult benchmarks for AI problem-solving and reasoning abilities.
OpenAI’s GPT-5 Pro has also self-reported perfect scores in both events.
Qwen3-Max-Thinking was built using Qwen3-Max, Alibaba’s largest AI model, which has over 1 trillion parameters and was unveiled in late September.
🔗 Source: South China Morning Post
🧠 Food for thought
Implications, context, and why it matters.
Benchmark wins lack verification
- Perfect scores on American Invitational Mathematics Examination (AIME) and Harvard-MIT Mathematics Tournament (HMMT) come from the vendor with no third-party check or public leaderboard.
- Missing details cover whether the test used 2025 problems, a closed-book run (no internet or tool access during testing), and contamination controls (checks to ensure test questions were not in training data).
- A two-week cryptocurrency trading test used Qwen3-Max’s 22.3% return versus GPT-5 Pro’s 62.7% loss, a window too short for statistical weight so luck can swamp skill.
- Independent, reproducible evidence with clear methods is still missing, so breakthrough claims for Qwen3-Max-Thinking on reasoning look early.
API access opens routing options, pricing and regions matter
- Software and data teams can try cost-performance routing by sending requests to Qwen3-Max-Thinking or other providers by price or accuracy to trim inference costs (the cost of generating model outputs) 1.
- Asia-Pacific (APAC) builders get a local option but still need API pricing plus regional availability beyond Singapore and quota limits with latency benchmarks for unit economics (per-request costs and margins) 12.
- Investors in AI infrastructure should watch adoption signals. If developers route complex reasoning through Qwen’s API while using Western providers for other work, that hints at competitiveness beyond benchmarks.
- Builders need batch inference availability (processing many requests to lower cost), context caching discounts (reusing previously processed tokens to avoid recomputation), and clarity on whether pay-per-token model (billing per unit of text) for multi-turn conversations (back-and-forth exchanges across messages) stays cost-effective versus rivals at scale 23.
Recent Alibaba developments
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.




