SuTRA tokenizer preserves word roots, boosting Indic MT by 8.08 chrF2
New algorithm fixes Morphological Shattering, improving translation quality by 8%
Standard subword tokenizers like BPE optimize for statistical compression, but they ignore morphological structure—a flaw that hurts morphologically rich Indic languages. These languages build words from aksharas (complex orthographic syllables), and frequency-based methods often over-fragment terms, arbitrarily splitting roots from affixes in a pattern the researchers call Morphological Shattering. To solve this, Vaibhav Rathore, Siddhant Gole, Dadhichi Telwadkar, Rooshil Bhatia, Maulik Ruparel, Siddharth Surekha, and Neha Bhargava propose SuTRA, an algorithm that treats aksharas as indivisible units and penalizes any merge that crosses a morphological boundary.
Accepted at Interspeech 2026, SuTRA shows dramatic gains over BPE: +14.7% in morphological alignment (Boundary F1) and +34% in semantic recoverability for Hindi. These structural improvements translate into real-world performance, yielding an average +8.08 chrF2 improvement in machine translation across tested language pairs. The authors also release a new morphological segmentation dataset for Hindi, Marathi, and Gujarati, giving researchers resources to advance morphology-aware tokenization further. SuTRA offers a practical path toward more accurate and interpretable NLP for Indic languages, with implications for translation, search, and language understanding at scale.
- SuTRA preserves akshara indivisibility and penalizes merges crossing morphological boundaries, preventing over-fragmentation
- Beat BPE by +14.7% in Boundary F1 and +34% in semantic recoverability for Hindi
- Improves machine translation by an average +8.08 chrF2; new segmentation dataset released for Hindi, Marathi, Gujarati
Why It Matters
Morphology-aware tokenization enables fairer, more accurate NLP for Indic languages, improving translation and search for over a billion speakers.