Engineering brief
Digital twins are 85% accurate—but only if you collect the right data
At a glance
- Relevance
- Practical value
- Warnings
- None
Simile AI claims 85% accuracy in predicting human behavior—3x better than frontier models on niche populations. The catch: they need bespoke interview and behavioral data, not just web text.
Human simulation could replace expensive panels and unlock causal decision-making at scale.
Summary
Joon Sung Park argues that the killer application for foundation models is human simulation, not personal assistants. His team achieved 85% accuracy in predicting how individuals behave in surveys and experiments, by combining in-depth interviews with behavioral data instead of relying only on web text. Frontier models alone scored 20-60% on niche populations.
The key claim is that current LLMs are super-rational and lack the messy, biased, and inefficient traits that make humans predictable. Simile AI collects data through consent-based interviews, transaction logs, and randomized control trials—what Park calls "dark knowledge"—to capture causal mechanisms rather than surface attitudes.
The tradeoff is significant: this approach requires bespoke data collection for each population, unlike combinatorial persona generation that scales cheaply but poorly. The vision is multi-agent simulations tackling wicked problems like climate change or democracy collapse, but current deployments focus on replacing human panels for market research and concept testing.
What's missing is independent validation of the 85% claim outside their own study, and clarity on how behavioral accuracy degrades when simulations interact in multi-agent settings. The cost argument hinges on saving expensive real-world experiments, not competing with cheap inference.
Watch the video
This video is blocked due to your privacy settings. To watch this video, please accept YouTube marketing cookies.
Related breakdowns
Protein design works. Scaling it is the hard part.
Chai Discovery's models now predict protein structures within an atom's width. The bottleneck is no longer model capability — it's integrating with pharma's…
The Real AI Moat Is an Assembly Line, Not an Algorithm
Poolside compressed frontier model training to 8 weeks via a 'model factory'—infrastructure speed is becoming AI's true moat.
Why the Next AI Scaling Axis Might Be a Wet Lab
Lila Sciences treats automated labs as verifiers, turning physical experiments into a scaling axis by generating tokens for generalist AI.
Get TL;DW
Too Long; Didn't Watch.
A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.
Free. Weekly. No hype.
Video and thumbnails remain the property of their respective creators. tldw.news provides editorial analysis, commentary, and discovery links to original content.