Research & Papers

Multilingual prompt injections can bypass defenses and inflate LLM relevance scores

Attackers can use 8 languages to trick LLM judges into ranking irrelevant documents higher.

Deep Dive

A new paper by Vo, Tuong, Zendel, and Sanderson explores how LLMs used as automated relevance judges can be manipulated via multilingual adversarial prompt injections. The authors used TREC Deep Learning datasets and two open-weight models (specific models not named) under standard prompting frameworks. They tested both instruction-based and content-based injection strategies across 8 languages spanning different resource levels, from high-resource (English, Spanish) to low-resource (Vietnamese, Tamil). The goal was to unethically inflate relevance scores for irrelevant documents while evading existing prompt-injection defenses.

The results showed that multilingual query-based injections were highly effective: they consistently inflated relevance judgments and successfully evaded state-of-the-art defense mechanisms. Even when defenses were modified to detect multilingual content, the injections were easily adapted to bypass these updates. The authors conclude that language generalization itself becomes an attack vector, as defenders cannot simply block low-resource languages without breaking functionality. This highlights a critical gap in current defense approaches and underscores the urgent need for proactive, language-aware evaluation frameworks for LLM-as-a-judge systems.

Key Points
  • Multilingual prompt injections inflated relevance scores across all 8 tested languages, regardless of resource level.
  • Existing prompt-injection defense mechanisms were consistently evaded by cross-lingual attacks.
  • When defenses were modified, attackers could easily adapt injections to bypass them, showing language as a attack surface.

Why It Matters

LLM-as-judge systems are vulnerable to multilingual attacks, demanding new robust evaluation frameworks.

📬 Get the top 10 AI stories daily