Why open-weights don’t fix AI’s cost problem
This article summarizes an episode of 20VC’s video series featuring Clay Bavor, co-founder of Sierra.

Clay Bavor, co-founder of Sierra/ Photo credit: Sierra
Clay Bavor, co-founder of Sierra, argues that open models will take over more routine work while the most capable systems stay valuable because they can handle higher-stakes tasks.
His broader point is that companies should stop treating AI as a simple software cost curve. They need to plan for limited access to advanced intelligence, larger token budgets, and agents that become part of the workforce.
Open models will handle more work while demand for smarter AI grows
Many AI buyers expect cheaper models to reduce their dependence on the largest AI labs.
Open-weights models, whose settings can be used and adapted by others, may fill that role for simple, routine work as they improve. But this does not answer what companies will do when stronger models can solve problems they could not automate before.
Bavor argues, “If you asked any software company, ‘Would you like to upgrade your staff-level software engineers to principal-level engineers?’ a hundred out of a hundred would say yes. We have not yet appreciated the unbounded demand for frontier levels of intelligence.”
Companies will need to choose the right model for each job. Basic customer support, search, and rule-based tasks can move to cheaper systems.
Coding, legal review, science, complex sales, and high-risk decisions may still need expensive model runs, because a better answer can be worth far more than the cost of running it.
Bavor also sees strong Chinese open models as part of a pricing fight. In his view, some of their strength comes from taking knowledge learned by closed models in the US and packaging it into open weights. Openness can be both a research choice and a way to pressure prices.
Reasoning makes tokens a real cost limit
If open models lower prices for routine tasks while demand grows for harder ones, the cost question changes. Companies cannot assume every model run will keep getting cheaper in a straight line.
The main cost driver is the extra computing used when a model reasons through a difficult problem. These models spend more tokens checking steps, comparing options, and trying again.
Hardware will improve, and older models will get cheaper, but harder questions may still require more computing power.
Bavor points to a hard physical constraint behind token pricing. “If you have unbounded demand for frontier-level intelligence… and the rate limiter is the number of Blackwells and H100s you have, you end up with a floor on the cost of tokens because you have to pay for the energy and the compute,” he warns.
Customer launches turn custom work into product knowledge
AI-first companies treat agents as part of the workforce
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.







