- Premium Content It takes our newsroom weeks - if not months - to investigate and produce stories for our premium content. You can’t find them anywhere else.
AI helps robots see in the real world, but blind spots persist
One morning, the engineers at Hand Plus Robotics set up their cameras to record videos of robots at work. It didn’t go as planned.
They were creating videos for visual language models (VLMs) – the robotic AI equivalent of a large language model (LLM) but built on images instead of just text. The robots use this visual data to learn how to pick things up, move them around, or even put them in a box.

Image credit: Tech in Asia
VLMs turn images and videos into words, which can be used to generate medical reports from x-rays or create product tags and visual descriptions. Visual language action (VLA) models go a step further, converting perception into action: where to move, how to grip, and how much force to apply.
The videos that morning looked clean, but when they tested the data on a robot, it “went blind.” The robot was no longer able to see its environment, according to Albert Causo, founder of Hand Plus Robotics.
The problem? Someone had left a window open. This let in strong morning sunlight that created a shadow over the robots during the video – a visual anomaly that ruined the model. The new data had to be tossed out.
This is a common issue and is also known as context or localization problems, says Causo.
His company does not build robots; it integrates and trains to be part of the client’s workflow. So the Singapore-based firm has learned to watch out for anything that can confuse the AI, the founder says.
Each environment – every camera angle, light source, and background – creates unique conditions, he tells Tech in Asia. Causo holds a Ph.D. in robotics.
Those difficulties highlight one of the toughest challenges in robotics today, namely teaching machines to see the world reliably. And this problem is just one of a number of issues causing a major choke point in the development of robotic AI: the lack of training data.
“Most of the time, there’s no data to work with,” Causo says. “The initial part of the conversation with clients always begins with what kind of data can be produced?”
Unlike text-based LLMs, which are fed by the internet’s seemingly endless reservoir of words, there is no widely accessible trove of data designed for robots.
“Unlike GPT, you can’t scrape Reddit or Wikipedia” to teach a robot, says Keechin Goh, founder of Singapore-based Datature, which manages datasets and fine-tunes vision models.
Knowing your robot’s brain
Transformer architecture – most commonly associated with LLMs – has fundamentally changed how robot “brains” are built. Developers are now moving away from narrow, hard-coded routines for these machines to VLMs and VLAs.
It’s not so simple
You can’t make this stuff up
Twist and shout
Stay ahead in Asia’s tech landscape
This is premium content. Subscribe to read the full story.
Robots can’t scrape the internet like GPT models – they need lots of real-world videos. That shortage is the biggest roadblock to smarter automation.
We know this is not ideal. ⌛ Sign up in 20 seconds. Cancel anytime.
Our subscriber community includes professionals from these companies:





Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.
