AI That Reads Emotions From Your Face Isn't as Good as Advertised
Combining face and voice barely boosts accuracy — and the inflated numbers may hide it.
A single-author study asks whether the usual "oracle ceiling" in audio-visual emotion recognition — the share of examples where at least one branch is right — is really a fusion target. It isn't, the paper argues. Using frozen self-supervised audio and visual encoders with a trained head, the author measured that ceiling and the realized gain on CREMA-D with two audio encoders (one whose fine-tuning lineage includes CREMA-D, one clean) and on EAV with the contaminated encoder. Swapping in a self-supervised encoder tripled the headroom, from +0.055 to +0.186. A second clean encoder and a controlled degradation of the contaminated one landed on the same curve, meaning provenance shifts headroom by changing audio-branch accuracy. Concatenation converted at most about half the headroom (0.55), and on EAV that converted fraction was indistinguishable from zero despite larger headroom. A learned router lost to plain concatenation everywhere. Oracle headroom, the paper concludes, measures branch disagreement, not a fusion budget.
- Mixing face and voice for emotion AI delivers only about half the accuracy people assume — and with a cleaned-up model, sometimes none at all.
- One audio model scored better simply because it had already seen the test data, showing how easily AI results get inflated.
- A smarter 'which signal should I trust?' system did worse than the simple method of just combining both, so added complexity isn't paying off.
Why It Matters
Emotion-reading AI is being sold for hiring, driving, and mental health — this study says its accuracy claims are inflated.