Engineering brief
The Inference Moat: How Cosine Builds Frontier AI on Millions
At a glance
- Relevance
- Practical value
- Warnings
- None
Cosine is building a UK sovereign LLM by skipping inference entirely—licensing model weights instead of hosting—to escape the unsolvable inference cost trap. If they succeed, the real AI moat shifts from training scale to inference overhead.
Reveals a blueprint for building competitive LLMs without matching Big Lab budgets by dodging inference costs and focusing on enterprise-aligned data.
Summary
Cosine, a UK coding-agent startup, won government backing to build a sovereign LLM using free Isambard supercomputer compute. They avoid the inference cost trap by licensing model weights, not hosting, sidestepping massive inference clusters and focusing resources on a narrow, enterprise-aligned model.
Pullen argues that architectural choices like active parameter count and dense models are critical; open-source labs lag because they optimize for inference efficiency over capability. Cosine plans a model above 100B active parameters, funded by tight feedback loops with large UK enterprises that shape training data—a data flywheel unavailable to generalized models.
To improve with less, Cosine experiments with credit attribution in RL: rewarding process, not just outcome. They pinpoint decision points in trajectories and adjust rewards to reduce “slop” and teach reusable abstractions, contrasting with correctness-only RL that produces spaghetti code.
The team’s agentic swarm automatically decomposes complex tasks, though Pullen believes harnesses will matter less over time. The interview surfaces a tension: sovereign AI is possible with narrow scope and clever cost avoidance, but scaling and reliability remain open questions.
Watch the video
This video is blocked due to your privacy settings. To watch this video, please accept YouTube marketing cookies.
Related breakdowns
Your AI agent harness is overengineered. The model got better.
Agents-as-files: Google DeepMind shows how markdown instructions replace Python agent loops. Cursor replaced 12,000 lines of TypeScript with 200 lines. But…
Durable execution is the real agent infrastructure challenge
Giselle van Dongen demonstrates why durable execution infrastructure, not agent SDKs, is the real bottleneck for production agent systems. Concrete failure…
AI agents need bank accounts and institutional memory, not better models
Coinbase treats every human correction to an AI PR as permanent repository memory. Agent payments and memory loops are becoming operational realities.
Get TL;DW
Too Long; Didn't Watch.
A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.
Free. Weekly. No hype.
Video and thumbnails remain the property of their respective creators. tldw.news provides editorial analysis, commentary, and discovery links to original content.