Research & Papers

MCBench: New benchmark exposes safety flaws in Omni LLMs

Omni LLMs can't integrate audio, vision, and text for safety

Deep Dive

Existing multimodal safety benchmarks only test visual inputs, ignoring the combined vision, audio, and text capabilities of Omni LLMs. To close this gap, Manh Luong and eight co-authors from various institutions introduced MCBench, a benchmark comprising 1,196 carefully constructed scenarios across four safety categories (e.g., physical harm, hate speech). Each unsafe scenario is paired with a minimally different safe counterpart to test model sensitivity—a design that forces models to rely on subtle multimodal cues rather than surface-level patterns.

Evaluations of state-of-the-art Omni LLMs revealed significant gaps. Models performed better when risks were accompanied by salient visual or acoustic cues but struggled with nuanced or non-physical dangers. Analysis of reasoning traces showed that while models could extract modality-specific information (e.g., recognizing a weapon in an image or a threatening tone in audio), they consistently failed to integrate these cues for accurate safety judgments. The findings underscore that current Omni LLMs lack robust cross-modal reasoning in safety-critical settings, pointing to an urgent need for improved architectures and training strategies that explicitly teach models to fuse information from multiple modalities before making safety decisions.

Key Points
  • MCBench includes 1,196 scenarios spanning four safety categories designed for Omni LLMs (multimodal models).
  • Each unsafe scenario is paired with a nearly identical safe counterpart to test model sensitivity to subtle cues.
  • Models can extract modality-specific info but fail to integrate vision, audio, and text for safety judgments.

Why It Matters

As multimodal AI enters critical applications, safety benchmarks must test cross-modal reasoning, not just single-modality.

📬 Get the top 10 AI stories daily