Engineering brief
Why AI agents work for code but fail elsewhere—and what to do
At a glance
- Relevance
- Practical value
- Warnings
- None
Karan Vaidya argues coding agents succeed because Git, CI/CD, tests, and governance built trust. Knowledge work lacks these six primitives—centralization, history, context, verification, governance, reversibility.
The bottleneck for AI agents shifts from models to missing infrastructure for non-coding domains.
Summary
Karan Vaidya argues that coding agents succeed because software engineering already has essential infrastructure: centralized repos, history (Git), testing frameworks, governance (code owners, branches), and reversibility. These systems enable trust without blocking speed.
Knowledge work agents lack all of this. A deal is scattered across Salesforce, Notion, Gmail, and Slack. No central source of truth, no history, no automatic verification. This forces agents to stitch context manually and act blindly, making even simple tasks dangerous.
The six primitives he proposes—centralization, record/history, context, verification, governance, and reversibility—are structural, not model-driven. Without them, agents cannot be trusted with sensitive actions like email or financial transactions.
While Composio is building this platform, the talk highlights a real tradeoff: building infrastructure is harder than improving models. Teams should focus on creating sandboxed environments and logging before deploying agents beyond code.
Watch the video
This video is blocked due to your privacy settings. To watch this video, please accept YouTube marketing cookies.
Related breakdowns
Multi-agent AI's real problem is privacy governance, not model power
Multi-agent AI faces a privacy governance bottleneck. The most practical approach: define a low-sensitivity zone where LLMs can make autonomous data-sharing…
Agent safety moves from models to runtime-level governance
Agent intelligence is almost solved. The real challenge is safely granting dynamic, scoped access at runtime. Docker’s new runtime aims to provide that, but…
Multiplayer agents demand team-wide infrastructure, not just better models
How to scale agentic coding across your team: shared sessions, secure sandboxes, and benchmarking on your own codebase—not just public benchmarks.
Get TL;DW
Too Long; Didn't Watch.
A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.
Free. Weekly. No hype.
Video and thumbnails remain the property of their respective creators. tldw.news provides editorial analysis, commentary, and discovery links to original content.