Engineering brief
Your AI Benchmarks Are Useless Without a Cost Axis
This engineering brief covers Your AI Benchmarks Are Useless Without a Cost Axis, with practical context for AI and developer-tool decisions.
The Brief
Noam Brown argues that AI benchmarks ignore inference budget, so a model's score can be bought with more compute rather than reflecting true capability. This misleads both model selection and safety governance.
Decision relevance
Read this for workflow impact, implementation trade-offs, and the claims that need technical scrutiny before they reach team planning.
Summary
Noam Brown says standard benchmark grids are broken because modern models improve greatly with more thinking time or compute, and the plateau is too distant to measure. Without an x-axis for cost, tokens, or time, a higher score could just come from more inference budget, and a better-looking model may be much more compute-efficient.
This distortion misleads both capability assessments and safety governance. Preparedness frameworks from the GPT-3 era don't account for test-time compute scaling. A model might pass low-budget dangerous-capability evaluations yet become harmful when an adversary spends more on inference.
The current release cycle (every 2–3 months) means nobody fully tests the ceiling before the next model arrives, leaving latent capabilities untested. The Erdos conjecture example showed that with a large enough budget ($10K–$100K of compute), earlier models could have achieved breakthroughs that weren't discovered until later.
Engineering leaders should stop relying on static benchmarks. Instead, define per-task budgets (latency, cost, tokens) and evaluate performance curves. This changes procurement, architecture, and safety governance. Beware routing or multi-agent claims that only look better due to more compute. Demand vendors publish curves, not isolated numbers, and build internal evals controlling for inference spend.
Why It Matters
Model capability is now a function of inference spend; static benchmarks mislead selection, budgeting, and safety governance.
Editorial analysis
Key claims
- Don't trust static benchmark grids; always ask: at what inference budget? Capability is a curve.
Practical use cases
- Use this as input for tooling evaluation, workflow planning, and technical due diligence.
Risks / caveats
- Overhyped multi-agent claims that don't control for inference budget may just be spending more compute.
Who should care
- Engineering managers, tech leads, and CTOs evaluating AI or developer tooling decisions.
Related topics
Bottom Line
Don't trust static benchmark grids; always ask: at what inference budget? Capability is a curve.
Watch
This video is blocked due to your privacy settings. To watch this video, please accept YouTube marketing cookies.
Related breakdowns
Moats Are Dead: AI Cuts Costs But Won’t Protect You
Booking’s AI is cutting costs, but its CEO warns there’s no moat. The real challenge? Proving ROI before scaling—and retraining teams before they’re displaced.
The RLHF Trap: Why AI Is Great at Chat but Terrible at
RLHF made AI great at conversation but terrible at automation. A former OpenAI researcher explains why, and what engineering leaders should do about it.
Dexterous manipulation is the bottleneck for general-purpose robots
Google DeepMind's Gemini Robotics 2 tackles dexterous manipulation—the unsolved bottleneck for general-purpose robots. Data scarcity, hardware limits, and…
Get TL;DW
Too Long; Didn't Watch.
A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.
Free. Weekly. No hype.
Video and thumbnails remain the property of their respective creators. tldw.news provides editorial analysis, commentary, and discovery links to original content.