🧔♂️ A friendly human may check it before it goes live. More news here
Tencent, Chinese university unveil new test for AI driving
🔍 In one sentence
Researchers introduced AD2-Bench, a benchmark designed to test the reasoning abilities of Multi-Modal Large Models (MLLMs) in autonomous driving tasks under adverse weather conditions.
🏛️ Paper by:
University of Chinese Academy of Sciences, Tencent CDG
✏️ Authors:
Zhaoyang Wei et al.
🧠 Key discovery
The study presents AD2-Bench as the first benchmark that evaluates Chain-of-Thought (CoT) reasoning in MLLMs specifically for autonomous driving in challenging weather scenarios. This fills a gap in current benchmarks, which typically overlook the complexity of real-world driving conditions.
📊 Surprising results
- Key stat: State-of-the-art models achieved less than 60% accuracy on AD2-Bench, showing the difficulty of the benchmark.
- Breakthrough: The dataset includes over 5,400 manually annotated CoT instances, enabling detailed analysis of reasoning in MLLMs.
- Comparison: Model performance on AD2-Bench is significantly lower than on earlier benchmarks, underlining the need for improved reasoning in adverse conditions.
📌 Why this matters
Current evaluation methods often fail to reflect the demands of real-world scenarios, such as heavy rain or fog. AD2-Bench provides a more realistic test environment, aiming to improve the interpretability and robustness of autonomous driving systems.
💡 What are the potential applications?
- Training improvements for MLLMs in adverse weather driving tasks.
2. Development of autonomous systems better suited to complex conditions.
3. Safer deployment of autonomous technologies in real-world settings.
⚠️ Limitations
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.




