- Insights This article was written by a TIA community member. Insights pieces undergo the same rigorous editorial process that newsroom-produced articles have.
The real money in AI isn’t in the models
Every week brings a new headline about which company has the smartest AI model. It’s the most exciting race in technology right now, but it’s the wrong bet for investors.
The competition over frontier models is a capital war among a handful of big players like OpenAI and Anthropic, and the prize keeps getting cheaper to win. But investors should pay attention to a different trophy: When the dust settles, which layer of the AI stack actually keeps the money?

Image credit: Ulla
Two numbers frame the answer. The cost of running a given level of AI capability has fallen by roughly 1,000x in three years, so the models are commoditizing fast. Over the same period, AI inference has grown to around US$118 billion in 2026 and now accounts for about two-thirds of all AI compute.
Intelligence is getting cheaper while the cost for delivering it keeps getting larger. That gap is where the opportunity lies for investors.
Becoming a commodity
Before going further, it’s worth pinning down definitions for two key terms.
Inference is what happens every time an AI model processes an input and produces an answer. This input can be a prompt being typed, a coding assistant running, or an agent firing off a task.
Model serving is the infrastructure behind these processes: the systems that route requests to the right hardware, manage cost and latency, and keep answering when millions of requests land at once.
These two elements are part of the AI value chain layer that this piece is about.
See also: The global hunt for AI compute has SEA in its crosshairs
The layer that gets the headlines – the AI models – is already being commoditized. That’s not merely a prediction; it is apparent in their prices.
GPT-4 launched in early 2023 at US$30 per million input tokens, meaning the foundational units of the text, code, or other data processed by an AI model. Nowadays, the cost of a model of equivalent capability is roughly 12x less, once better and more affordable successors are taken into account.
Research institute Epoch AI put the decline for equal performance at around 10x per year. VC firm A16z made the same estimate and even has a name for it: LLMflation.
Instead of offering a temporary discount, strong open-weight models such as Llama, DeepSeek, Qwen, and Mixtral keep closing the performance gap with closed-source frontier models. Several labs now ship near-parity systems, and a combination of distillation (a smaller model training on the output of a more developed one) and better hardware grinds down the cost every quarter.
Value in the middle
The middle layer is scaling
No free lunch
What it means
Stay ahead in Asia’s tech landscape
This is premium content. Subscribe to read the full story.
As AI models commoditize, the reliable money is moving to the infrastructure layer that serves them – a growing market valued at US$118 billion.
We know this is not ideal. ⌛ Sign up in 20 seconds. Cancel anytime.
Our subscriber community includes professionals from these companies:





Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.

