Research & Papers

Qwen Small Models Ignore Conflicting Instructions, Study Shows

Even accurate small models routinely disobey when instructions clash with task behavior.

Deep Dive

A new arXiv paper from Mahdiyeh Farajidizaji and Vatsal Raina investigates whether small language models (SLMs) follow instructions when those instructions conflict with the model's learned task behavior. The researchers tested three tasks: multiple-choice question answering (MCQA), sentiment classification, and mathematical QA. For each, they paired a standard instruction (e.g., "select the correct answer") with a conflicting non-standard instruction (e.g., "select an incorrect option," "output the opposite sentiment," or "return twice the answer"). Using Qwen models of varying sizes, they measured standard accuracy, non-standard accuracy, and an Instruction-Following Failure Rate (IFFR).

Results show that small models remain competent on standard tasks but routinely ignore the conflicting instruction, appearing accurate under standard metrics. Larger models exhibit a clear gap between the two settings, indicating that instruction following does not automatically improve with scale. The authors argue that task capability and instruction adherence are separate abilities, and reporting only standard accuracy masks instruction-following failures. This has implications for deploying SLMs in contexts where precise instruction following is critical.

Key Points
  • Evaluated Qwen models (various sizes) on MCQA, sentiment, and math tasks with conflicting instructions.
  • Introduced IFFR (Instruction-Following Failure Rate) to quantify when models ignore non-standard instructions.
  • Small models maintain task accuracy but often fail to follow conflicting instructions; larger models show a clear accuracy gap.

Why It Matters

Reliable instruction following is critical for deployment; standard accuracy alone is misleading in small models.

📬 Get the top 10 AI stories daily