Topic
Developer Tooling - Page 5
Tools that change how teams build, review, and ship. Curated tldw.news briefings about developer tooling, with practical engineering takeaways from long-form AI and developer-tool videos.
159
breakdowns
Page 5 of 16
AI EngineerRealistic Evals, Not Benchmark Scores, Gate Your AI Agents
Lyft’s team shows how realistic offline simulation and validated, actionable LLM judges create a production safety net for AI agents, not just a score.
AI EngineerGraph Design Cut AI Tool Calls 40% in Code Search
A 40% drop in AI code-search tool calls wasn't from better models—it came from graph algorithms. But only if you build the graph right.
AI EngineerAI Model Quality Is Table Stakes—UX Now Decides Adoption
Many AI features fail because users don't trust them. This talk breaks down the UX patterns—citations, transparency, control—that can make or break adoption.
Theo - t3․ggFable 5 vs GPT-5.6: The Real Cost Is Merge Debt
Two top AI coding models, two opposite tradeoffs: cost vs merge quality. Fable 5 is the premium plan, Soul the fast executor. Here's when to use each.
AI EngineerLoops Won’t Replace Your Engineers—Unless You Let Them
Loops promise speed but demand discipline: verification gaps, cost, and code rot threaten to undercut the hype. Pragmatic teams start small.
David OndrejManaging AI Agents, Not Just Using Them
An AI engineer uses an agent manager and adversarial code review catching 63% of AI-generated changes—a shift in how teams should design AI workflows.
IBM TechnologyThe Model Wars Are Over—Customization Is the New Moat
Inkling’s open-weight fine-tuning platform vs. Muse Spark’s agent focus: The model race is no longer about benchmarks but who gives teams the most control.
AI EngineerExecution Is Cheap—Your Eval Design Will Make or Break Success
An AI agent won a coding competition by executing human ideas, signaling execution is automated. High-leverage work shifts to evaluation design, architecture.
Theo - t3․ggYour Model Is Fine, Your System Prompt Is Sabotaging You
Codex’s hidden system prompt mandates specific border radii and bans empty states, wasting tokens and producing generic output. Fix the harness, not the model.
InfoQCutting SDK Duplication in Half with a Rust Core
Temporal's shared Rust core halved SDK duplication, but the bridge layer's complexity is the real lesson for leaders scaling multi-language products.
Get TL;DW
Too Long; Didn't Watch.
A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.
Free. Weekly. No hype.