Build Once, Probe Many Times
The paper’s key move is to turn each participant into a reusable agent scaffold, then test that scaffold on hidden tasks rather than training a new predictor for every downstream question.
Instead of prompting a generic demographic persona, this paper builds reusable LLM agents from what specific people actually said in interviews, surveys, or both, then asks whether those agents generalize to unseen questions and experiments.
The paper’s key move is to turn each participant into a reusable agent scaffold, then test that scaffold on hidden tasks rather than training a new predictor for every downstream question.
The normalization only makes sense because people do not answer identically every time. This 1,052-cell grid visualizes an illustrative retest-consistency distribution behind the human ceiling idea.
Coverage and stereotype pressure move in opposite directions. Richer self-report packets expose more person-specific signal while reducing reliance on demographic shorthand.
Interviews contribute narrative constraints, surveys contribute standardized coverage, and the combined agent blends both into one reusable prompt-and-memory scaffold.
The episode gives exact toplines for the four agent types. Family-level splits below are realistic reconstructions used to show how the pattern likely looks across evaluation types.
The transcript highlights smaller accuracy gaps across racial and ideological groups once the agent knows what a person said instead of only who they resemble demographically.
Illustrative deltas are scaled to the episode’s qualitative claim: self-report grounding narrows group gaps relative to the demographics-only baseline.
The closer a task stays to language-based self-reports, the safer the paper’s claim looks. Push farther toward revealed behavior, temporal drift, or high-stakes decisions and the confidence should collapse fast.
The episode’s practical line is narrow but clear: pretest questions and simulate heterogeneous respondent pools, but do not treat synthetic people as substitutes in hiring, lending, benefits, or clinical decisions.
Key papers from the episode lineage and critique arc. Transcript-mentioned arXiv IDs do not add beyond the known source paper, so the cards below focus on the cited background and caution literature.
Charts marked as reconstructed or illustrative use the transcript’s exact toplines and generate plausible splits only to make the visual comparison legible.