Engineering brief
Better data is the cheapest compute multiplier you're ignoring
This engineering brief covers Better data is the cheapest compute multiplier you're ignoring, with practical context for AI and developer-tool decisions.
The Brief
Compute is getting scarce and expensive. But Ari Morcos shows that data curation can deliver 100x efficiency gains.
Decision relevance
Read this for workflow impact, implementation trade-offs, and the claims that need technical scrutiny before they reach team planning.
Summary
Compute scarcity is worsening fast. H100 prices rose 40%, reasoning models use 8x more tokens, and API access is being capped. In this environment, Ari Morcos argues that data quality is the overlooked lever: better data steepens the learning curve, delivering the same performance for far less compute.
DatologyAI's results show a 14-point gain in VLM benchmarks through curation alone, matching Qwen 3 72B with 145x less training compute. For text models, curated multilingual data improved non-English MMLU performance using only 8% multilingual tokens. The claim is supported by multiple controlled comparisons, not just marketing slides.
However, the talk is a company pitch. The core insight—data curation changes scaling laws—is real, but DatologyAI positions itself as the necessary tool. The practical question for teams is whether internal curation pipelines can achieve similar gains without vendor lock-in.
For engineering leaders, the actionable signal is clear: invest in data infrastructure, not just more compute. But be skeptical of vendor claims that only their approach works. Open-source curation methods are advancing fast.
Why It Matters
Data curation can reduce training costs by 100x, making specialized model training viable for more teams.
Editorial analysis
Key claims
- Invest in data curation infrastructure before buying more compute.
Practical use cases
- Use this as input for tooling evaluation, workflow planning, and technical due diligence.
Risks / caveats
- The subtle product pitch for DatologyAI's platform.
Who should care
- Engineering managers, tech leads, and CTOs evaluating AI or developer tooling decisions.
Related topics
Bottom Line
Invest in data curation infrastructure before buying more compute.
Watch
This video is blocked due to your privacy settings. To watch this video, please accept YouTube marketing cookies.
Related breakdowns
Your LLM Bill Is a Model Selection Problem
Most features don't need frontier models; evaluate SLMs with a golden dataset and prompt engineering to match quality and eliminate inference costs.
The Real Scaling Problem for AI Delivery Isn’t Autonomy—It’s Operations
DoorDash’s AI ordering boosts discovery, but their in-house delivery robot reveals hidden ops challenges. Plus, a 20x AI spend spike that forced ROI discipline.
Napkin Math Exposes the Real Cost of AI Infrastructure
Turbopuffer’s napkin math reveals 100x cost gaps in vector search. Learn how first-principles thinking can transform AI infrastructure spend.
Get TL;DW
Too Long; Didn't Watch.
A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.
Free. Weekly. No hype.
Video and thumbnails remain the property of their respective creators. tldw.news provides editorial analysis, commentary, and discovery links to original content.