Engineering brief

Your AI Benchmarks Are Useless Without a Cost Axis

No Priors: AI, Machine Learning, Tech, & Startups2 min read · saves 34 min

At a glance

Relevance
Practical value
Warnings
None

Noam Brown argues that AI benchmarks ignore inference budget, so a model's score can be bought with more compute rather than reflecting true capability. This misleads both model selection and safety governance.

Model capability is now a function of inference spend; static benchmarks mislead selection, budgeting, and safety governance.

Summary

Noam Brown says standard benchmark grids are broken because modern models improve greatly with more thinking time or compute, and the plateau is too distant to measure. Without an x-axis for cost, tokens, or time, a higher score could just come from more inference budget, and a better-looking model may be much more compute-efficient.

This distortion misleads both capability assessments and safety governance. Preparedness frameworks from the GPT-3 era don't account for test-time compute scaling. A model might pass low-budget dangerous-capability evaluations yet become harmful when an adversary spends more on inference.

The current release cycle (every 2–3 months) means nobody fully tests the ceiling before the next model arrives, leaving latent capabilities untested. The Erdos conjecture example showed that with a large enough budget ($10K–$100K of compute), earlier models could have achieved breakthroughs that weren't discovered until later.

Engineering leaders should stop relying on static benchmarks. Instead, define per-task budgets (latency, cost, tokens) and evaluate performance curves. This changes procurement, architecture, and safety governance. Beware routing or multi-agent claims that only look better due to more compute. Demand vendors publish curves, not isolated numbers, and build internal evals controlling for inference spend.

Watch the video

This video is blocked due to your privacy settings. To watch this video, please accept YouTube marketing cookies.

Related breakdowns

Get TL;DW

Too Long; Didn't Watch.

A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.

Free. Weekly. No hype.

Video and thumbnails remain the property of their respective creators. tldw.news provides editorial analysis, commentary, and discovery links to original content.