AI audits online hate across languages: Phi-4 detects shared moral grammar
Phi-4 achieves 95% accuracy in labeling hate across English, Chinese, and Malay posts from Singapore.
A new study by Emilio Ferrara cross-lingually audits online hate in multicultural Singapore using eight open LLMs as annotators. Analyzing 31 million social media items and 1.76 million comments from Facebook, Reddit, and YouTube (2025), the work identifies Phi-4 as the best performing model, achieving 0.95 accuracy, 0.91 Cohen's κ, and 1.00 recall against a human-adjudicated gold set. The core finding—termed 'layered cultural contingency'—shows that while which out-groups are targeted varies significantly by language (Cramer's V=0.25), the threat frames and moral grammar of hate are highly shared: sanctity and loyalty dominate (55–75%) over fairness, with divergence dropping to V=0.08 for moral foundations and 0.07 for emotion.
Hate is predominantly contempt-driven, anti-immigrant (not anti-system), and reception is selectively nativist: hateful comments are amplified less than neutral mentions overall, yet anti-immigrant hate is preferentially amplified while religious and anti-LGBTQ hate is not. Volume does not track key Singapore events in 2025. Critically, the study shows absolute hate prevalence is unreliable across LLM annotators (agreement ceilings at κ≈0.42), so relative structure is reported as primary. These results directly inform cross-lingual content moderation strategies, highlighting the need for culturally aware, structurally focused moderation tools.
- Phi-4 achieves 95% accuracy and 100% recall in hate detection across English, Chinese, and Malay posts.
- Anti-immigrant hate is preferentially amplified; religious and anti-LGBTQ hate is not, despite similar volumes.
- Moral foundations of hate (sanctity, loyalty) are 88% shared across languages; only targets are culturally specific (V=0.25).
Why It Matters
This study reveals that hate speech moderation must focus on shared moral structures, not just target identities, to work across languages.