Tokyo Researcher's AI Runs a Million 'People' in One Second
Cheaper than a survey, faster than a trial — but can you trust it?
Social scientists, governments and marketers constantly want the same thing: to know how people will react to a new policy, ad or public message before they commit to it. The usual answer is a survey or a field trial with real humans — slow, expensive and sometimes impossible. KITE, a new system from the University of Tokyo, offers a shortcut: instead of studying real people, it creates large crowds of AI stand-ins, called agents (AI programs that make decisions on their own), and watches what they do.
The clever part is the cost. Asking a top-tier AI model about every single simulated person would be painfully slow and pricey. So KITE asks the expensive model only in a small number of situations — about 1.7% of the cases in one test — and uses those answers to correct the rest, which run from a pre-built table. In experiments modeled on a study of 9,070 real participants, that small sample cut the error in predicted effects by 41%. Across 37 held-out social science experiments, it improved how well the simulation matched real decisions, from 0.27 to 0.39. And a million agents finished 20 rounds of decisions in 0.9 seconds on an ordinary laptop.
Just as important is what the tool admits it doesn't know. Most AI simulations act confident. KITE measures the gap between its AI stand-ins and real humans, then carries that uncertainty into every conclusion it draws. That honesty showed up in testing: when it claimed 80% or 90% confidence, it was right about 93% and 96% of the time, versus only 29% and 36% when relying on human sampling noise alone. It also passed content checks in all 15 new countries added to a 16-country study.
The catch: these are still make-believe people. The agents follow simplified decision rules, and the system was validated against studies that already existed, not against fresh real-world predictions. Treat KITE as a cheap first filter — a way to screen which ideas are worth testing on actual humans — not as a replacement for them.
- KITE lets researchers test policies and messages on large crowds of AI stand-ins for real people instead of running costly human trials.
- A million simulated people made 20 decisions in under a second on a laptop, and checking just 1.7% of cases with a powerful AI cut prediction error by 41%.
- It reports its own uncertainty honestly, but the 'people' are simplified rules — a screening tool, not a substitute for real studies.
Why It Matters
Governments and companies could test policies on fake crowds first, saving money and avoiding bad real-world rollouts.