Tired of ads? Enjoy an ad-free experience by signing up.
Grace Priscilla Teo · · 5 min read

The hidden problem with multi-step AI workflows

This article summarizes an episode of EO’s video series  featuring Abhishek Das, co-founder and co-CEO of Yutori.

Image credit: Ulla

Many automated tools fail, and users often accept it. Abhishek Das, co-founder and co-CEO of Yutori, sees this as a fatal flaw. He argues that this low standard forces business leaders to rethink reliability, software design, and trust in web automation.

Fixing this low standard starts with understanding exactly why these tools break down.

The math behind tasks with many steps

Software teams often release automated tools that look good during quick tests. They launch products packed with features and simply rely on user patience when things inevitably break.

When engineers look at high success rates for single steps, they miss how those numbers drop when steps are chained together.

“If we think of a [multi-step workflow], a 10-step, 20-step, or 50-step workflow,” Das argues, “even if each step is 90% accurate, the 10% error rate compounds quickly, making overall success rate quite low.”

This mathematical reality destroys big AI startup promises. Das notes, “There are basically 100 different agent products saying that [an agent] can do anything on the web, and you try it once and it doesn’t really work.”

He refuses to accept this low standard: “If it’s not good enough to work on the first try, it’s not good enough.” Building lasting automation requires strict rules, not just shipping early bugs.

Moving through changing websites

Building reliable AI systems needs engineering effort beyond writing simple text prompts. Treating software as a series of decisions in a chaotic, changing internet means that finding and fixing errors must become the primary goal.

Because the internet is too vast to prepare for every page layout, systems must handle surprises safely. Das believes reliable agents agents need to detect their own mistakes, reverse actions, and pivot when they hit a dead end.

He emphasizes, “Safely stopping and recalculating is always cheaper than executing a blind mistake.”

Mandating exhaustive quality checks
To achieve this safe recalculation, engineering teams must add safety checks. “We put in a lot of effort into building [tests] and guardrails,” Das notes.

“Every single production query that a user runs goes through a fairly comprehensive set of tests that lets us quickly identify where these agents are doing well versus not,” he says.

Finding the problems users ignore

Checking quality through internal testing

The business value of showing your work

Building trust through design



Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.

TIA Writer

Grace Priscilla Teo

A Singapore-based writer with a passion for AI, cats, and donuts. Grace covers emerging tech and AI developments, bringing fresh insights with a uniquely personal touch. (AI-generated profile.)