OpenAI's GPT-generated surveys rival human-designed ones in social science study
LLM-crafted surveys capture same attitudinal divisions as expert instruments across three domains.
A new study from Tina Behzad, Wenbo Li, Reuben Kline, and Klaus Mueller evaluates whether large language models like GPT can replace human experts in designing social attitude surveys. The researchers created GPT-generated surveys using a fixed prompting framework that enforced a 3x3 structure over beliefs, perceptions, and behaviors across three domains: climate change, immigration, and diversity, equity, and inclusion (DEI). These were compared to validated human-designed baselines matched for length and construct coverage. U.S.-based participants completed both survey types, enabling direct within-subject comparison.
The results show that GPT-generated surveys capture the same dominant attitudinal divisions as human-designed instruments, but exhibit differences in the resolution of belief structure and group separation. The authors conclude that LLM-generated surveys are well-suited for exploratory and large-scale analyses, and can complement expert-designed instruments. This finding could accelerate survey design for social science research, making large-scale attitude measurement more accessible and scalable.
- GPT-generated surveys were compared to human baselines across climate change, immigration, and DEI using a within-subject design with US participants.
- GPT surveys captured the same dominant attitudinal divisions as validated human surveys but with different resolution in belief structure.
- The study suggests LLM-generated surveys are effective for exploratory and large-scale analyses, complementing expert-designed instruments.
Why It Matters
LLMs could automate survey generation for large-scale social research, saving time and resources while maintaining reliability.