Research & Papers

New Arabic-Russian LLM benchmark breaks language barriers for science

Qwen2.5 fine-tuned on 27K sentence pairs, boosting BLEU by +4.36

Deep Dive

M. K. Arabov introduces a new benchmark to bridge the scientific communication gap between Arabic and Russian researchers. The work presents a hybrid parallel corpus of about 27,000 sentence pairs, drawn from scientific abstracts as well as general-domain texts like religion, news, and conversations. The benchmark fine-tunes three multilingual language models—mT5-base (580M parameters), NLLB-200-distilled-1.3B (1.3B), and Qwen2.5-7B-Instruct (7B)—using LoRA with ranks 8, 16, 32, and 64.

The best results came from Qwen2.5-7B with QLoRA at rank 8, yielding BLEU 23.15, chrF 43.89, BERTScore 0.906, and COMET 0.758. This represents a +4.36 BLEU and +0.051 COMET improvement over the zero-shot baseline. Few-shot prompting with three examples did not improve performance, indicating that domain-specific fine-tuning is necessary for meaningful scientific translation. The researcher has released all models, the corpus, and evaluation code. This work directly supports UN Sustainable Development Goals 9 (industry, innovation, and infrastructure) and 17 (partnerships for the goals) by lowering language barriers in research.

Key Points
  • 27K sentence-pair hybrid corpus of scientific abstracts and general-domain texts (Arabic-Russian).
  • Qwen2.5-7B fine-tuned with QLoRA (rank 8) achieves +4.36 BLEU over zero-shot baseline.
  • Few-shot prompting failed; domain-specific fine-tuning is required for scientific translation.

Why It Matters

Enables Arabic and Russian scientists to exchange findings, accelerating global collaboration and sustainable development.

📬 Get the top 10 AI stories daily