Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

OpenAI’s new reasoning models see rise in hallucination rates

OpenAI’s latest AI models, o3 and o4-mini, show higher hallucination rates compared to previous versions.

Internal tests found improvements in coding and math, but also more inaccurate or fabricated responses.

The o3 model hallucinated 33% of the time in the PersonQA benchmark, up from 16% in o1 and 14.8% in o3-mini. The o4-mini model scored even higher, with a 48% hallucination rate.

Third-party testing by Transluce found similar issues. The o3 model, for instance, falsely claimed it had run code on a 2021 MacBook Pro outside the ChatGPT environment.

OpenAI acknowledged these challenges and is exploring solutions, including adding web search capabilities.

🔗 Source: TechCrunch


🧠 Food for thought

1️⃣ OpenAI’s regression reverses a multi-year trend of improving accuracy

The increased hallucination rates in OpenAI’s o3 and o4-mini models represent a significant departure from historical patterns in AI development.

Recent benchmark studies show hallucination rates typically decreased with newer models, from 40% in ChatGPT 3.5 to 29% in ChatGPT 4 according to UX Tigers research 1.

This improvement trend was consistent, with the Hugging Face Hallucination Leaderboard documenting a regression of approximately 3 percentage points per year in hallucination rates. However, predictions that AI could potentially reach zero hallucinations by 2027 1 may have been overly optimistic given the reversal with o3, which hallucinates 33% of the time on PersonQA, double the rate of previous reasoning models. This disrupts the pattern of steady improvement and suggests that scaling reasoning capabilities may introduce new challenges for factual reliability.

These findings align with academic research, where a systematic review of reference generation found significant differences between models. For instance, GPT-4 had a 28.6% hallucination rate compared to GPT-3.5’s 39.6% 2.

2️⃣ The business cost of AI hallucinations extends beyond simple errors

OpenAI’s hallucination challenges highlight a persistent problem that companies across industries are actively struggling to address as they implement AI.

Businesses report that AI hallucinations lead to significant operational disruptions, regulatory compliance issues, and reputational damage, particularly in sectors where accuracy is critical such as healthcare, finance, and legal services 3.

Recent surveys indicate 77% of businesses express serious concerns about the accuracy of their AI systems, recognizing that hallucinations can result in legal liability, customer dissatisfaction, and operational inefficiencies 4.

Recent OpenAI developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.