Tired of ads? Enjoy an ad-free experience by signing up.
  • Premium Content
    It takes our newsroom weeks - if not months - to investigate and produce stories for our premium content. You can’t find them anywhere else.
Scott Shuey · · 7 min read

AI helps robots see in the real world, but blind spots persist

One morning, the engineers at Hand Plus Robotics set up their cameras to record videos of robots at work. It didn’t go as planned.

They were creating videos for visual language models (VLMs) – the robotic AI equivalent of a large language model (LLM) but built on images instead of just text. The robots use this visual data to learn how to pick things up, move them around, or even put them in a box.

Image credit: Tech in Asia

VLMs turn images and videos into words, which can be used to generate medical reports from x-rays or create product tags and visual descriptions. Visual language action (VLA) models go a step further, converting perception into action: where to move, how to grip, and how much force to apply.

The videos that morning looked clean, but when they tested the data on a robot, it “went blind.” The robot was no longer able to see its environment, according to Albert Causo, founder of Hand Plus Robotics.

The problem? Someone had left a window open. This let in strong morning sunlight that created a shadow over the robots during the video – a visual anomaly that ruined the model. The new data had to be tossed out.

This is a common issue and is also known as context or localization problems, says Causo.

His company does not build robots; it integrates and trains to be part of the client’s workflow. So the Singapore-based firm has learned to watch out for anything that can confuse the AI, the founder says.

Each environment – every camera angle, light source, and background – creates unique conditions, he tells Tech in Asia. Causo holds a Ph.D. in robotics.

Those difficulties highlight one of the toughest challenges in robotics today, namely teaching machines to see the world reliably. And this problem is just one of a number of issues causing a major choke point in the development of robotic AI: the lack of training data.

“Most of the time, there’s no data to work with,” Causo says. “The initial part of the conversation with clients always begins with what kind of data can be produced?”

Unlike text-based LLMs, which are fed by the internet’s seemingly endless reservoir of words, there is no widely accessible trove of data designed for robots.

“Unlike GPT, you can’t scrape Reddit or Wikipedia” to teach a robot, says Keechin Goh, founder of Singapore-based Datature, which manages datasets and fine-tunes vision models.

Knowing your robot’s brain

Transformer architecture – most commonly associated with LLMs – has fundamentally changed how robot “brains” are built. Developers are now moving away from narrow, hard-coded routines for these machines to VLMs and VLAs.

It’s not so simple

You can’t make this stuff up

Twist and shout

Stay ahead in Asia’s tech landscape

This is premium content. Subscribe to read the full story.

Why subscribe?

Robots can’t scrape the internet like GPT models – they need lots of real-world videos. That shortage is the biggest roadblock to smarter automation.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

10

10 company database access

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

🧠 For professionals / ⭐ Best value

CoreBest value

US$16.58US$14.92/month

Billed annually at US$179.10 on the first year

Get instant access to this article and more every month

Unlimited premium content

Unlimited news briefs & articles

Unlimited company database access

Ad-free reading experience

Just US$0.55 per day

Save US$19.90 on the first year. Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.

TIA Writer

Scott Shuey

Scott has worked as a journalist for over 20 years, including 18 years working in Asia. He covers emerging technologies such as AI and Web3. You can reach him at scott.shuey@techinasia.