Research & Papers

Scientists Can't Tell Why People Text Back — And That Skews Outbreak Forecasts

Models that predict disease and viral posts may be off by 2.5x

Deep Dive

Scientists often map how things spread using "temporal networks" — simply put, records of who contacted whom and when. There are two competing explanations for why contact comes in bursts. Either certain people are repeat initiators (they just message a lot), or specific relationships are strengthening through repeated contact. The new paper proves that when you can see who started each contact, these two explanations are cleanly separable. When you can't — as with phone proximity data, where two devices just register as near each other — the two get tangled together and stay tangled even after sophisticated statistical cleanup.

Testing real proximity, messaging, and email records, however, the author found something blunter and worse: the "chatty person" effect was essentially stuck to a simple pattern in the timing of events, barely budging no matter how the model was adjusted. A shuffle test — randomly reordering events to destroy any real memory — showed the model was recovering a chatty-person signal from nothing at all. Refitting with far more flexible math didn't help. The same collapse showed up in bot-edited Wikipedia-style networks, cloud server logs, and even brain cell firing data.

The practical fallout is about epidemics. When researchers simulated outbreaks using these fitted contact records, they under-predicted how many people would get infected by up to a factor of 2.5, and the tipping point at which an outbreak takes off shifted. That means public health planners using this style of model could size a response wrongly — too few vaccines, too little warning, or the reverse.

The honest limitation: recording who initiates a contact genuinely fixes half the problem, but heavy-tailed "chatty person" behavior appears impossible to identify from timing alone. The lesson for anyone relying on spread predictions — from disease teams to social media analysts — is to treat them as rougher estimates than they look.

Key Points
  • Contact data can't reliably tell 'this person messages everyone' apart from 'these two are growing closer' — two very different explanations for the same pattern
  • On real phone, text, and email records, the model's guess barely moved, and a shuffle test showed it 'found' memory in randomly scrambled data
  • Epidemic simulations built on that data underestimated outbreak size by up to 2.5 times and shifted the tipping point where outbreaks take off

Why It Matters

Outbreak forecasts and viral-content predictions built on contact data may be significantly wrong, affecting public health decisions.

📬 Get the top 10 AI stories daily