Engineering brief
Your AI Benchmarks Are Useless Without a Cost Axis
At a glance
- Relevance
- Practical value
- Warnings
- None
Noam Brown argues that AI benchmarks ignore inference budget, so a model's score can be bought with more compute rather than reflecting true capability. This misleads both model selection and safety governance.
Model capability is now a function of inference spend; static benchmarks mislead selection, budgeting, and safety governance.
Summary
Noam Brown says standard benchmark grids are broken because modern models improve greatly with more thinking time or compute, and the plateau is too distant to measure. Without an x-axis for cost, tokens, or time, a higher score could just come from more inference budget, and a better-looking model may be much more compute-efficient.
This distortion misleads both capability assessments and safety governance. Preparedness frameworks from the GPT-3 era don't account for test-time compute scaling. A model might pass low-budget dangerous-capability evaluations yet become harmful when an adversary spends more on inference.
The current release cycle (every 2–3 months) means nobody fully tests the ceiling before the next model arrives, leaving latent capabilities untested. The Erdos conjecture example showed that with a large enough budget ($10K–$100K of compute), earlier models could have achieved breakthroughs that weren't discovered until later.
Engineering leaders should stop relying on static benchmarks. Instead, define per-task budgets (latency, cost, tokens) and evaluate performance curves. This changes procurement, architecture, and safety governance. Beware routing or multi-agent claims that only look better due to more compute. Demand vendors publish curves, not isolated numbers, and build internal evals controlling for inference spend.
Watch the video
This video is blocked due to your privacy settings. To watch this video, please accept YouTube marketing cookies.
Related breakdowns
Arm CEO: AI verification is the real chip design bottleneck
Arm CEO: AI is revolutionizing chip verification, but supply chain and data center constraints are the next hurdles.
Moats Are Dead: AI Cuts Costs But Won’t Protect You
Booking’s AI is cutting costs, but its CEO warns there’s no moat. The real challenge? Proving ROI before scaling—and retraining teams before they’re displaced.
Anthropic's safety layering creates hidden non-determinism for agent workflows
Anthropic's safety-layered models create hidden non-determinism when classifiers silently swap engine behavior. The OpenAI Hugging Face escape shows…
Get TL;DW
Too Long; Didn't Watch.
A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.
Free. Weekly. No hype.
Video and thumbnails remain the property of their respective creators. tldw.news provides editorial analysis, commentary, and discovery links to original content.