Why the public AI progress wall is completely false to OpenAI
This article summarizes an episode of OpenAI’s video series featuring its research lead, Tejal Patwardhan.

Photo credit: Shutterstock
Evaluating technical progress requires looking beyond public benchmarks to measure genuine utility.
OpenAI research lead Tejal Patwardhan contends that organizations fixate too heavily on leaderboard scores, urging leaders to deploy realistic evaluations that track speed, safety, and readiness well before a product reaches the public.
Models learn skills long before people use them
Most executives judge AI by daily interactions, assuming it isn’t ready when they spot an awkward email or factual mistake. Early software often has rough edges that hide how fast the underlying system improves.
Testing teams operate on a different schedule, running models in secure laboratory environments to discover what the software can accomplish right now.
“Evals are a way to measure and understand what models can do,” Patwardhan explains, “and see progress before it tends to happen.”
Confusing a lack of public adoption with technical weakness misguides company planning. The actual question for leaders centers on whether an AI system can perform a specific task given the right laboratory tools and permissions.
Proper testing functions as a critical early warning system, showing exactly which jobs computers will soon handle so teams can prepare for changes before signing new contracts or finalizing annual budgets.
Continuous improvement beats temporary flaws
Technical development outpaces public adoption, which changes how organizations should plan. The next challenge is measuring that pace, especially since many users look at current mistakes and assume the software has already peaked.
Internal research tells a different story. Newer systems show continuous progress as they spend more time calculating answers to solve harder problems in subjects like chemistry and physics.
“Hitting the wall is just so not the right way to think about model progress,” Patwardhan argues. Having tracked these systems for an extended period, she notes the technology keeps improving, with no signs of slowing on the current research roadmap.
This steady progress breaks traditional purchasing cycles: a six-month security review can leave the tested software outdated before the pilot even finishes.
Executives need to track how fast tools improve at real daily tasks, testing new versions against internal goals instead of relying on stale metrics.
Custom tests work better than public rankings
Independent software programs break traditional tests
Testing in the physical world demands new skills
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.





