LLM synthetic population test passes all 7 controllability criteria
120 fictional personas exposed to institutional messages reveal internal validity across temperature sweeps.
Mirko Degli Esposti's paper, 'Calibrating the Instrument: Controllability of an LLM-Driven Synthetic Population,' tackles a fundamental question before generative synthetic populations (GSPs) can be used for real-world urban simulations: do they respond to stimuli in an ordered, replicable, group-structured way? The author calls this property 'controllability' and argues it must be established before any claim of external validity (matching real humans). To measure it, he designed SIVE (Synthetic Instrument Validation Experiment), a simulated municipality called Montelago with 120 synthetic personas of known latent structure.
The experiment exposed these personas to seven communication conditions about a water network, ranging from strongly positive to strongly negative. Seven pre-registered criteria—fidelity, stability, noise floor, specificity, sensitivity, and ordering—were evaluated across a full temperature sweep. Remarkably, all seven criteria passed at every temperature setting. A standout finding turned a calibration failure into a diagnostic success: a message designed as 'weakly positive' was identified by the synthetic population as functionally negative. Tracing this back to the text revealed unresolved problems, uncertainty, and institutional passivity in the wording. A redesigned version restored the expected ordering and even interacted with agents' latent trust in unexpected ways.
Further analysis included a noise sub-experiment showing that the instrument's intrinsic noise was roughly half the cross-agent estimate and stable across temperatures. Individual agent trajectories revealed coherent micro-dynamics that summary statistics would obscure. All data and an interactive explorer are available. This work provides a rigorous framework for validating LLM-driven simulations before deployment.
- All 7 pre-registered criteria (fidelity, stability, noise floor, specificity, sensitivity, ordering) passed across a full temperature sweep with 120 synthetic personas.
- A 'weakly positive' message was actually perceived as negative by the LLM agents, traced to text containing unresolved problems and institutional passivity.
- Intrinsic noise of the synthetic population was roughly half the cross-agent estimate and remained stable across temperatures.
Why It Matters
Ensures LLM-driven synthetic populations are reliable for urban policy simulations before external validation is attempted.