AI Learns That You're the Same Person Across All Your Data
Your apps could finally stop confusing you with someone else.
Machine learning is usually formalized through samples—but a new paper argues that misses the persistent referent those samples come from. The paper proposes making the “unit” an explicit primitive: a learning task declares a population of persistent referents and a sameness criterion. In the supervised setting, what gets learned is a tokenizer that produces a contextual unit token, plus a shared response-law form that reads it. A key formal result: unlinked single-row observations can fail to distinguish a heterogeneous unit world from a homogeneous pooled world, while trusted same-unit pairs can separate a restricted witness. It’s a rigorously grounded case for putting the individual—defined as a unit, not a person—back at the center of learning.
- Most AI today treats each data point as independent, ignoring that many points come from the same person.
- The paper introduces a 'unit token' so the model can carry identity through its predictions.
- When identity is unknown, the system can infer it from clues, but a single unlinked data point can't reveal whether people are truly different.
Why It Matters
This could lead to AI that personalizes better, makes smarter medical decisions, and never mixes you up with a stranger.