Tired of ads? Enjoy an ad-free experience by signing up.
  • Insights
    This article was written by a TIA community member. Insights pieces undergo the same rigorous editorial process that newsroom-produced articles have.
Kyle Qi · · 5 min read

The real money in AI isn’t in the models

Every week brings a new headline about which company has the smartest AI model. It’s the most exciting race in technology right now, but it’s the wrong bet for investors.

The competition over frontier models is a capital war among a handful of big players like OpenAI and Anthropic, and the prize keeps getting cheaper to win. But investors should pay attention to a different trophy: When the dust settles, which layer of the AI stack actually keeps the money?

Image credit: Ulla

Two numbers frame the answer. The cost of running a given level of AI capability has fallen by roughly 1,000x in three years, so the models are commoditizing fast. Over the same period, AI inference has grown to around US$118 billion in 2026 and now accounts for about two-thirds of all AI compute.

Intelligence is getting cheaper while the cost for delivering it keeps getting larger. That gap is where the opportunity lies for investors.

Becoming a commodity

Before going further, it’s worth pinning down definitions for two key terms.

Inference is what happens every time an AI model processes an input and produces an answer. This input can be a prompt being typed, a coding assistant running, or an agent firing off a task.

Model serving is the infrastructure behind these processes: the systems that route requests to the right hardware, manage cost and latency, and keep answering when millions of requests land at once.

These two elements are part of the AI value chain layer that this piece is about.

See also: The global hunt for AI compute has SEA in its crosshairs

The layer that gets the headlines – the AI models – is already being commoditized. That’s not merely a prediction; it is apparent in their prices.

GPT-4 launched in early 2023 at US$30 per million input tokens, meaning the foundational units of the text, code, or other data processed by an AI model. Nowadays, the cost of a model of equivalent capability is roughly 12x less, once better and more affordable successors are taken into account.

Research institute Epoch AI put the decline for equal performance at around 10x per year. VC firm A16z made the same estimate and even has a name for it: LLMflation.

Instead of offering a temporary discount, strong open-weight models such as Llama, DeepSeek, Qwen, and Mixtral keep closing the performance gap with closed-source frontier models. Several labs now ship near-parity systems, and a combination of distillation (a smaller model training on the output of a more developed one) and better hardware grinds down the cost every quarter.

Value in the middle

The middle layer is scaling

No free lunch

What it means

Stay ahead in Asia’s tech landscape

This is premium content. Subscribe to read the full story.

Why subscribe?

As AI models commoditize, the reliable money is moving to the infrastructure layer that serves them – a growing market valued at US$118 billion.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

10

10 company database access

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

🧠 For professionals / ⭐ Best value

CoreBest value

US$16.58US$14.92/month

Billed annually at US$179.10 on the first year

Get instant access to this article and more every month

Unlimited premium content

Unlimited news briefs & articles

Unlimited company database access

Ad-free reading experience

Just US$0.55 per day

Save US$19.90 on the first year. Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.

Community Writer

Kyle Qi

Kyle Qi is a principal at Silicon Valley–based early-stage VC firm Llama Ventures.