Engineering brief
Mission-Critical Generative AI in Action • Scott Shaw • YOW! 2025
At a glance
- Relevance
- Practical value
- Warnings
- None
Scaling GenAI into production requires a platform with gateway, guardrails, and evaluation, plus a shift to evaluation-driven development.
It provides a battle-tested blueprint for engineering leaders to move GenAI from pilot to production safely and efficiently at scale.
Summary
Scott Shaw, from CommBank, shares hard-won lessons from running GenAI at massive scale (80B tokens/week). The core problem: most prototypes stay experimental because teams hit unpredictable latency, crushing inference costs, frequent model deprecation, and safety concerns that stall production. His solution is a minimal viable platform built on three pillars—a central gateway that provides one API, load balancing, and auditing; tuned guardrails that filter inputs and outputs against safety and compliance needs; and an evaluation framework that collects, labels, and statistically compares model performance.
Without this platform, teams repeatedly spend months revalidating models every time a frontier model reaches end-of-life (often less than a year). The talk also redefines engineering practice: instead of test-driven development, adopt evaluation-driven development with statistical assertions, and instead of starting with the most powerful model, pick the smallest, cheapest model that meets quantitative thresholds. This requires data pipelines, synthetic bootstrapping, and continuous human-in-the-loop supervision.
For leaders, the signal is clear: enthusiastic experimentation must be channeled into disciplined processes with platform support, or the gap between prototype and production remains a grave of wasted investment. The shift demands that every engineer develop model rigor, a risk mindset, and mathematical literacy in probability and statistics.
Watch the video
This video is blocked due to your privacy settings. To watch this video, please accept YouTube marketing cookies.
Related breakdowns
AI Governance Is the Real Bottleneck, Not Model Capability
AI isn't just a productivity tool—it's a governance challenge. The 2040 scenario shows why pacing and transparency matter more than raw capability.
Agents face the same operational debt as microservices—prepare now
Navan shares hard lessons from running agents in production: runtime is solved, but cost, testing, and debugging gaps threaten every team scaling agentic AI…
Stripe and IBM bet the model war is already over—routing is the
Stripe's $7B OpenRouter acquisition signals that routing, not models, is where AI value is moving. IBM's dual partnerships with OpenAI and Anthropic…
Get TL;DW
Too Long; Didn't Watch.
A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.
Free. Weekly. No hype.
Video and thumbnails remain the property of their respective creators. tldw.news provides editorial analysis, commentary, and discovery links to original content.