AHA-Memes benchmark brings 5K annotated Arabic memes to hate detection
Researchers release AHA-Memes, a 5K-meme benchmark for Arabic hateful meme detection.
Hateful memes are a growing form of multimodal online harm, where intent emerges from images, text, and cultural subtext. While English and other high-resource languages have seen advances in detection, Arabic remains far behind—existing resources mostly target propaganda or coarse labels. To close that gap, a team of researchers introduced AHA-Memes (Arabic HAteful Memes), a new benchmark built from 5K manually annotated memes using a fine-grained taxonomy of hate types and attack strategies. The team also released ~66K silver-labeled memes for future training, making it the first large-scale resource of its kind for Arabic.
To establish baselines, the researchers benchmarked text-only, image-only, and late-fusion multimodal models, as well as few-shot in-context learning and both open- and closed-weight vision-language models (VLMs) under zero-shot and fine-tuning settings. Results show that culturally grounded references—such as regional stereotypes and dialect-specific idioms—are especially challenging for current models, underscoring the need for dedicated Arabic datasets. The dataset, annotation guidelines, and evaluation scripts are publicly available, offering a foundation for safer social media moderation across Arabic-speaking communities.
- AHA-Memes provides 5K manually annotated Arabic memes with multi-label hate taxonomy
- Includes ~66K silver-labeled memes for training and future studies
- Benchmarks show vision-language models struggle with culturally grounded hate in zero-shot settings
Why It Matters
First large-scale Arabic hateful meme benchmark enables culturally aware moderation for over 400M Arabic speakers.