Engineering brief
Open models are winning on economics. The bottleneck is now orchestration and
At a glance
- Relevance
- Practical value
- Warnings
- None
Open models now power 150x token growth this year, driven by coding agents. The cost argument is winning—but the real prize is control and customization.
Cost is the wedge; control and customization are the long-term moat. Coding agents are the killer app.
Summary
Open models are shifting from hobbyist curiosity to enterprise infrastructure, driven by cost pressure and the rise of coding agents. Ollama's data shows token usage has grown 150x this year, with the surge coming from open models powering agentic workflows. Chinese models dominate cloud deployments, while US and European models compete locally. Cost reduction is
the entry point, but customization and control are the long-term goals. The real unlock is a new class of 'flash' models—good enough for 80% of tasks, ultra-cheap, and fast. These enable unlimited-token usage patterns, similar to early ChatGPT. The bottleneck has moved from model intelligence to coordination, security, and workflow design. Enterprises like AT&T have
already shifted 40% of token consumption to open models, primarily for coding agents. The key tradeoff is between frontier closed models for hard problems and open models for volume work. Most enterprises will use a hybrid model—80-90% of tokens from open models, but only 10-20% of budget. The scarcity is shifting from tokens to the
layers above: orchestration, memory, credential management, and safety tooling. Open models lack the out-of-box safety infrastructure of closed providers, creating both risk and opportunity for new tooling businesses. The geopolitical dimension is real but manageable. The 'Manchurian Candidate' concern exists theoretically, but enterprise security teams are experienced with open-source supply chain risks. Model provenance matters
Watch the video
This video is blocked due to your privacy settings. To watch this video, please accept YouTube marketing cookies.
Related breakdowns
Harness, not model, now decides agent performance and cost
How much agent performance comes from the harness, not the model? YC's harness night shows huge gains—and the governance costs that follow.
Leya's $100M lesson: Skip fine-tuning, bet on better models
Leya's $100M ARR in 18 months without fine-tuning any model. The real lesson: bet on model improvement, obsess over workflow design, and freeze sales when…
White House AI strategy: open source, preemption, and the science ecosystem shift.
The White House AI strategy doubles down on open source and preemption, but the real signal is the shift in science funding and the push for autonomous labs.
Get TL;DW
Too Long; Didn't Watch.
A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.
Free. Weekly. No hype.
Video and thumbnails remain the property of their respective creators. tldw.news provides editorial analysis, commentary, and discovery links to original content.