New AI Tool Spots Misleading Photo Captions in Nepali
Fake news with real photos tricks people—this AI catches it.
Researchers introduced NepOOC, the first publicly available Nepali-dominant multilingual benchmark for detecting out-of-context misinformation, where real images are paired with misleading captions. It contains 1,090 image-caption pairs (545 pristine, 545 out-of-context) annotated across five misinformation typologies with strong annotator agreement (kappa = 0.84). After testing five multimodal architectures alongside text-only and image-only baselines, they found a text-only mBERT model achieved 94.65±0.20% Macro-F1, statistically tied with the best multimodal system (ResNet-50+mBERT). Image-only models performed near chance (33–50%), and dataset expansion appeared to be a more direct route to progress than more sophisticated architectures.
- NepOOC is the first public dataset for detecting out-of-context misinformation in Nepali and English, with 1,090 image-caption pairs.
- A text-only AI model matched the best image-and-text models at 94.65% accuracy, while image-only models were near random.
- The findings suggest that for fighting misinformation in low-resource languages, expanding datasets is a more direct path than building complex AI systems.
Why It Matters
This research paves the way for better fake news detection in Nepali, potentially protecting millions from misleading social media posts.