- Premium Content It takes our newsroom weeks - if not months - to investigate and produce stories for our premium content. You can’t find them anywhere else.
Om AI bets on unusual combo: real videos and robot brains
While much of the AI industry races to conjure new videos from “nothing” – i.e., without filming – Zhao “Tony” Tiancheng is selling the opposite.
His goal is to build models that understand real footage rather than generate synthetic clips. “Real video still has big value in the future,” says Zhao, who is also a principal researcher at Zhejiang University’s Binjiang Institute.
The understanding works in two ways, giving Hangzhou Om AI Technology an unusual business: It is both an AI video company and a robotics company.

Zhao “Tony” Tiancheng, founder of Hangzhou Om AI / Photo credit: Hangzhou Om AI
Founded in 2021, Om AI has deliberately stayed out of the cloud-based models that seem to dominate current headlines. Instead, it is developing AI models that are small enough to run on PCs, cameras, and robots.
These models run at 1 billion to 10 billion parameters, not the trillion-scale of the largest systems, according to Zhao.
Rather than generate videos, Om AI’s model organizes and annotates real footage with place names, categories, and other metadata, making a video catalog searchable for users. Zhao argues that authentic videos still have value: A brand selling a phone, for example, wants buyers to see the real product, “not a fake phone generated by some AI.”
See also: The startups powering Asia’s AI-generated video boom
Users pay for a one-time hardware fee rather than a subscription. Zhao puts it at about US$15,000 for an enterprise team of roughly 10. Om AI is pushing for the “edge approach” as it is cheaper to run, lighter on data uploads, and stronger on privacy.
At Beyond Expo 2026 in Macao last month, Om AI launched OttoBox AI Studio, a video tool that runs on laptops and other small devices. The software allows users to analyze hundreds of gigabytes of footage – captioning clips, tagging objects and people, and making the whole library searchable – without paying for cloud processing, Zhao says.
He points out that the bottom line here is cost. Uploading raw footage to a cloud service and paying per token to process hours of unedited video wastes money on material that never makes the final cut. But with OttoBox handling that locally, video creators can use their cloud tokens for the polished output instead.
Zhao, who worked on early vision-language models while taking his doctorate in computer science at Carnegie Mellon University’s Language Technologies Institute, admits that one problem still eludes him. On the media side, the technology can describe what a video shows, but it cannot judge whether a shot is beautiful – a subjective call that he reckons only “movie masters” can make.
Om AI built OttoBox by working with professional broadcasters and large brands first, Zhao says. It then packaged what it learned into a standardized product aimed at small businesses and individual creators.
Not one but two
The access to video has led Om AI to create a “second” business. The company is using video to create and train a “brain box” – a visual language action model that can be plugged into existing robots, giving them the ability to move and follow spoken instructions.
Training over profits
Stay ahead in Asia’s tech landscape
This is premium content. Subscribe to read the full story.
The Hangzhou-based startup is building lightweight vision models for creators and robots, steering clear of the generative-video race.
We know this is not ideal. ⌛ Sign up in 20 seconds. Cancel anytime.
Our subscriber community includes professionals from these companies:





Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.
