Engineering brief
The AI safety net that nods along while patients are harmed
At a glance
- Relevance
- Practical value
- Warnings
- None
Production AI scribes have a 1 in 20 serious error rate. The best automated judges miss most of them because the hard part isn't catching differences, it's knowing which differences matter.
Silent AI errors in production now affect patient safety, and evaluation systems miss most of them.
Summary
A real-world study of production AI clinical scribes finds 1 in 20 notes contain errors serious enough to cause significant harm. These are not obvious hallucinations, but subtle mistakes: a headache note omits jaw pain on chewing, which would have revealed giant cell arteritis, a blindness emergency. The note is factually correct, but
dangerously incomplete. The core problem is taste: knowing what matters in a specific context. A dropped line about a patient's holiday is noise in one case and the diagnosis in another. Current evaluation systems fail because they rely on pre-specified rubrics that cannot capture this contextual judgment. The best automated judges pass
one in five seriously flawed notes. Verification is not easier than generation for these problems. Unlike code where unit tests exist, there is no free verifier for safety or completeness. The hard part is not spotting differences between transcript and note, but judging which differences matter. This requires tacit, contextual, and moving
knowledge that cannot be written down in advance. The solution is a continuous evaluation loop: discover failure modes from real outputs, capture expert judgment as examples, and calibrate each output against similar past cases. This approach outperforms static rubrics and fine-tuned models because it learns and adapts to evolving standards without retraining.
Watch the video
This video is blocked due to your privacy settings. To watch this video, please accept YouTube marketing cookies.
Related breakdowns
Agent value can’t be measured by tokens — build verification first
If you can’t measure agent output, you can’t justify the spend. Use verification difficulty—not token cost—to choose AI agent use cases.
AI made speed free. Signal and trust are the new moats.
AI made speed free for everyone. The new competitive advantage is signal—your specific, unaverageable point of view that AI can't replicate. Lena Hall…
Figma's blueprint for AI agent adoption without shipping garbage
Figma's engineering org reveals that AI agent adoption fails on culture and workflow, not technology. Their fix: planning-first development, deterministic…
Get TL;DW
Too Long; Didn't Watch.
A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.
Free. Weekly. No hype.
Video and thumbnails remain the property of their respective creators. tldw.news provides editorial analysis, commentary, and discovery links to original content.