AI Post Transformers / Visual Companion

Simulating Individuals with Self-Reported LLM Agents

Instead of prompting a generic demographic persona, this paper builds reusable LLM agents from what specific people actually said in interviews, surveys, or both, then asks whether those agents generalize to unseen questions and experiments.

Joon Sung Park et al. / 2024 / Stanford, Northwestern, Washington, Google DeepMind, Sciences Po. The episode centers on a 1,052-person holdout test: survey items, personality traits, behavioral games, and randomized interventions.

arXiv 2411.10109 1,052 participants 4 holdout families Combined agent: 86% Demographics-only: 74%
Best condition
86%
Interview plus survey agents approach a person’s own two-week answer consistency.
Weak baseline
74%
Thin demographic prompting trails richer person-specific evidence by a wide margin.
Core mechanism
Holdout
Targets are hidden before the agent is built, so evaluation probes generalization instead of recall.
Hard question
Language != life
Strong results on self-report tasks still do not prove deep behavioral simulation in the wild.

Build Once, Probe Many Times

Step through the pipeline

The paper’s key move is to turn each participant into a reusable agent scaffold, then test that scaffold on hidden tasks rather than training a new predictor for every downstream question.

Focus

Noisy Humans Are the Ceiling

Hover the cohort grid

The normalization only makes sense because people do not answer identically every time. This 1,052-cell grid visualizes an illustrative retest-consistency distribution behind the human ceiling idea.

Hover

Selected References

Key papers from the episode lineage and critique arc. Transcript-mentioned arXiv IDs do not add beyond the known source paper, so the cards below focus on the cited background and caution literature.

Source paper ID: 2411.10109

Charts marked as reconstructed or illustrative use the transcript’s exact toplines and generate plausible splits only to make the visual comparison legible.