MatrAIx simulates 8.3B personas to evaluate AI systems and digital products
8.3B simulated users, 1,290 personality traits, 91.5% accuracy — AI testing goes massive.
MatrAIx, introduced by a team of 93 researchers including Xiaomin Li, tackles a fundamental bottleneck in AI evaluation: human testing is slow, costly, and hard to scale. Their solution is a population-scale infrastructure built on Persona 8B, a dataset of 8.3 billion persona records defined by 1,290 categorical dimensions. These records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. The researchers release a quality-filtered coreset of roughly 1 million personas (599,847 human-grounded and 400,000 synthetic), providing a ready-made pool of heterogeneous simulated users.
The MatrAIx Playground offers four distinct environments — Survey, AI Chatbot, Web, and App — where persona agents evaluate digital products across 1,010 application tasks spanning commerce, software, finance, and healthcare. The team ran 18,189 evaluation trials across eight representative tasks, powering agents with Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The results captured nuanced behavior shifts: hesitation after price increases, reduced willingness to continue after an AI assistant fails, and varying latency tolerance by persona background. Validation was strong: a 400-trial controlled study found declared behavior was correctly expressed or suppressed in 91.5% of trials, while both human and LLM judges confirmed the quality of human-grounded persona extraction. This makes MatrAIx an end-to-end alternative to traditional user studies for AI system evaluation.
- Persona 8B covers 8.3 billion personas across 1,290 categorical dimensions, with a 1M-persona coreset released for public use.
- 18,189 evaluation trials ran across 4 environments and 1,010 tasks, powered by Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5.
- Persona adherence hit 91.5% in a 400-trial controlled validation study across 10 behavioral attributes.
Why It Matters
Replaces costly human testing by simulating millions of diverse users, enabling faster, more reliable AI product iteration at scale.