Research & Papers

New AI Dataset Aims to Fix Clunky Translation for Indian Languages

⚡Better translations for Hindi, Tamil and Bengali — without English in the middle.

Deep Dive

A team of 25 researchers from Indian universities and research labs has built COILD, a new public dataset designed to make AI translation work better between Indian languages. It contains 1.16 million sentences, each written by a human and checked by a human, covering 20 language pairs from four major language families — Indo-Aryan (like Hindi and Bengali), Dravidian (like Tamil and Telugu), Tibeto-Burman, and Austro-Asiatic. Crucially, the sentences were collected from original Indian-language sources across eight everyday domains, not translated out of English first.

That last detail is the whole point. Most AI translation tools today take a shortcut: to turn Hindi into Tamil, they quietly go Hindi to English to Tamil. It's like playing telephone with a middleman who keeps mishearing things. Idioms, honorifics, and local context get flattened. The new dataset gives AI models training material written naturally in each language, so translations can sound like what a real speaker would say. That matters for anything where getting words wrong has consequences: government forms, medical instructions, legal documents, school materials, and customer support for hundreds of millions of people.

The researchers also built a test set of 2,000 sentences, hand-verified by language experts, so anyone can fairly compare translation tools across Indian languages. To prove the data is useful, they trained two existing AI translation models on it — IndicTrans2-Distilled and NLLB-200 — and both got measurably better, according to automatic scoring and human reviewers. That's a signal the bottleneck wasn't smarter algorithms, it was better data.

The honest catch: this is a research resource, not an app you can download. It only covers 20 language pairs, while India has 22 official languages and hundreds of spoken ones. It's text-only, so no voice or video. And better data doesn't automatically reach your phone — companies still have to adopt it. Think of it as high-quality ingredients delivered to the kitchen, not a finished meal.

Key Points
  • COILD contains 1.16 million human-written and human-checked sentences across 20 Indian language pairs — a rare resource, since most AI translation data is built by translating from English first.
  • It fixes the 'English middleman' problem, where translating Hindi to Tamil actually goes Hindi → English → Tamil and loses meaning along the way.
  • Researchers also released a 2,000-sentence expert-verified test set and showed two existing AI translation models improved when trained on the new data.

Why It Matters

More accurate translation for hundreds of millions of people in official forms, healthcare, legal documents and daily messages — without an English detour.

📬 Get the top 10 AI stories daily