Study Reconstruction Pipeline
The model is not fed “a persona.” It is fed standardized tuples: respondent, condition, question, answer.
This page treats the episode like a benchmark dashboard: how 210 TESS studies become SOCSCI210, how fine-tuning changes the simulator, and where distribution fit still diverges from causal faithfulness.
Exact corpus counts come from the episode transcript. Performance plots use metric-faithful mock values shaped to the reported relative gains, so the visuals show structure, tradeoffs, and failure modes rather than pretending to reproduce the paper’s tables verbatim.
The model is not fed “a persona.” It is fed standardized tuples: respondent, condition, question, answer.
Each row is a tiny experimental scene. Fine-tuning teaches the model to map that scene onto an answer distribution, not just generate plausible prose.
Hover cells to see what gets preserved when messy study materials are flattened into training rows.
Hot cells mark studies that are harder to standardize or simulate: more arms, more outcome scales, or denser response volume.
Orders of magnitude matter here: participant counts, condition counts, and outcome counts do not live on the same scale.
The win condition is not “sounds human.” It is matching how human answers spread across bins and how that spread shifts under treatment.
The episode’s core warning lives in this quadrant: better distribution matching can still leave the simulator below the bar for causal trust.
One chart asks whether treatment direction survives. The other asks whether aggregate improvement still leaves hot subgroup disparities.
This companion lists the papers from the episode sources that have direct arXiv IDs.