Research & Papers

BaFCo benchmark exposes MLLMs' struggles with Bangla forms

New benchmark reveals major gaps in AI's ability to read complex Bengali documents.

Deep Dive

Document comprehension remains a tough challenge for Multimodal Large Language Models (MLLMs), especially for low-resource languages like Bangla. To address the lack of annotated data, researchers from multiple institutions introduced BaFCo, a benchmark dataset for Bangla form comprehension. BaFCo curates 200 multi-page complex government forms from sectors like agriculture, education, banking, and land management. The dataset features a fine-grained annotation schema with 26 entity types, plus a coarse set of 5 types, enabling detailed evaluation of Document Layout Analysis (DLA) and Key Information Extraction (KIE) tasks.

The team evaluated the latest MLLMs from the ChatGPT, Gemini, Claude, Qwen, and Kimi series using zero-shot and chain-of-thought prompts under both low and high reasoning setups. Results exposed clear limitations: all models struggled to accurately localize highly granular form entities in Bangla. The work, accepted at ECCV 2026, provides a crucial benchmark for advancing multilingual AI. The dataset and code are publicly available, offering a foundation for future research in low-resource language document understanding.

Key Points
  • Dataset includes 200 multi-page Bangladeshi government forms from 5 sectors (agriculture, education, banking, land management)
  • Fine-grained annotation schema with 26 entity types and a coarse set of 5 types for DLA and KIE
  • Evaluated 5 MLLM series (ChatGPT, Gemini, Claude, Qwen, Kimi) showed significant localization failures

Why It Matters

Highlights critical gaps in AI document understanding for low-resource languages, limiting global adoption of MLLMs.

📬 Get the top 10 AI stories daily