Engineering brief

Better data is the cheapest compute multiplier you're ignoring

AI Engineer1 min read · saves 18 min

At a glance

Relevance
Practical value
Warnings
None

Compute is getting scarce and expensive. But Ari Morcos shows that data curation can deliver 100x efficiency gains.

Data curation can reduce training costs by 100x, making specialized model training viable for more teams.

Summary

Compute scarcity is worsening fast. H100 prices rose 40%, reasoning models use 8x more tokens, and API access is being capped. In this environment, Ari Morcos argues that data quality is the overlooked lever: better data steepens the learning curve, delivering the same performance for far less compute.

DatologyAI's results show a 14-point gain in VLM benchmarks through curation alone, matching Qwen 3 72B with 145x less training compute. For text models, curated multilingual data improved non-English MMLU performance using only 8% multilingual tokens. The claim is supported by multiple controlled comparisons, not just marketing slides.

However, the talk is a company pitch. The core insight—data curation changes scaling laws—is real, but DatologyAI positions itself as the necessary tool. The practical question for teams is whether internal curation pipelines can achieve similar gains without vendor lock-in.

For engineering leaders, the actionable signal is clear: invest in data infrastructure, not just more compute. But be skeptical of vendor claims that only their approach works. Open-source curation methods are advancing fast.

Watch the video

This video is blocked due to your privacy settings. To watch this video, please accept YouTube marketing cookies.

Related breakdowns

Get TL;DW

Too Long; Didn't Watch.

A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.

Free. Weekly. No hype.

Video and thumbnails remain the property of their respective creators. tldw.news provides editorial analysis, commentary, and discovery links to original content.