The hidden problem with multi-step AI workflows
This article summarizes an episode of EO’s video series featuring Abhishek Das, co-founder and co-CEO of Yutori.

Image credit: Ulla
Many automated tools fail, and users often accept it. Abhishek Das, co-founder and co-CEO of Yutori, sees this as a fatal flaw. He argues that this low standard forces business leaders to rethink reliability, software design, and trust in web automation.
Fixing this low standard starts with understanding exactly why these tools break down.
The math behind tasks with many steps
Software teams often release automated tools that look good during quick tests. They launch products packed with features and simply rely on user patience when things inevitably break.
When engineers look at high success rates for single steps, they miss how those numbers drop when steps are chained together.
“If we think of a [multi-step workflow], a 10-step, 20-step, or 50-step workflow,” Das argues, “even if each step is 90% accurate, the 10% error rate compounds quickly, making overall success rate quite low.”
This mathematical reality destroys big AI startup promises. Das notes, “There are basically 100 different agent products saying that [an agent] can do anything on the web, and you try it once and it doesn’t really work.”
He refuses to accept this low standard: “If it’s not good enough to work on the first try, it’s not good enough.” Building lasting automation requires strict rules, not just shipping early bugs.
Moving through changing websites
Building reliable AI systems needs engineering effort beyond writing simple text prompts. Treating software as a series of decisions in a chaotic, changing internet means that finding and fixing errors must become the primary goal.
Because the internet is too vast to prepare for every page layout, systems must handle surprises safely. Das believes reliable agents agents need to detect their own mistakes, reverse actions, and pivot when they hit a dead end.
He emphasizes, “Safely stopping and recalculating is always cheaper than executing a blind mistake.”
Mandating exhaustive quality checks
To achieve this safe recalculation, engineering teams must add safety checks. “We put in a lot of effort into building [tests] and guardrails,” Das notes.
“Every single production query that a user runs goes through a fairly comprehensive set of tests that lets us quickly identify where these agents are doing well versus not,” he says.
Finding the problems users ignore
Checking quality through internal testing
The business value of showing your work
Building trust through design
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.





