Engineering brief
Napkin Math Exposes the Real Cost of AI Infrastructure
At a glance
- Relevance
- Practical value
- Warnings
- None
Simon Ericson's napkin math at Shopify exposed misleading database benchmarks. His startup Turbopuffer uses S3 caching to cut vector search costs by 100x, proving first-principles analysis reveals massive infrastructure savings.
Napkin math reveals 100x cost mismatches in vector search; simplicity and first-principles thinking can beat vendor hype for engineering leaders.
Summary
Simon Ericson’s "napkin math" approach—calculating theoretical hardware limits—exposed that many database benchmarks are misleading. At Shopify, he used it to challenge infrastructure decisions, leading to tools like Toxyroxy for failure testing. The lesson: engineering leaders should demand first-principles cost and performance estimates before adopting tools.
After leaving Shopify, Ericson applied napkin math to vector search, realizing S3’s durability could slash costs if latency could be managed. His prototype used simple clustering and Nginx caching, offering $1 per million vectors—undercutting incumbents by 100x. This forced rethinking where AI infrastructure spend goes: from expensive managed services to commodity storage.
Cursor, desperate to fix unit economics, became the first customer after Ericson personally debugged their Postgres issues. The 95% cost reduction proved that reliability and simplicity can win over feature-heavy competitors. The risk: betting on a one-person startup paid off, but only because of deep technical trust built through hands-on help.
Now, Turbopuffer faces new constraints: CPU and NVMe shortages as RL and agent workloads explode. Ericson’s CPU-centric architecture benefits from broad SKU compatibility, but the fight for compute is real. The strategic takeaway: infrastructure bets must account for shifting resource scarcity, not just current pricing.
Watch the video
This video is blocked due to your privacy settings. To watch this video, please accept YouTube marketing cookies.
Related breakdowns
AI Governance Is the Real Bottleneck, Not Model Capability
AI isn't just a productivity tool—it's a governance challenge. The 2040 scenario shows why pacing and transparency matter more than raw capability.
Agents face the same operational debt as microservices—prepare now
Navan shares hard lessons from running agents in production: runtime is solved, but cost, testing, and debugging gaps threaten every team scaling agentic AI…
Stripe and IBM bet the model war is already over—routing is the
Stripe's $7B OpenRouter acquisition signals that routing, not models, is where AI value is moving. IBM's dual partnerships with OpenAI and Anthropic…
Get TL;DW
Too Long; Didn't Watch.
A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.
Free. Weekly. No hype.
Video and thumbnails remain the property of their respective creators. tldw.news provides editorial analysis, commentary, and discovery links to original content.