Research & Papers

C4 benchmark reveals MLLMs fail at creative cross-concept reasoning

Best closed models score only 50.7% on decoding creative Chinese idioms...

Deep Dive

A new arXiv paper introduces C4, a cognition-inspired framework for evaluating cross-concept creativity using Chengyu (Chinese idioms). The C4-Eval set includes 184 synthetic items and 37 human-created figures, instantiated across 5 task settings for 884 cases. Testing 10 MLLMs, the strongest closed models reached only 50.7% and 48.0% accuracy, while open-source models lagged substantially behind. Candidate constraints improved accuracy sharply, but bridge hints and explanation requests gave only modest gains — exposing a major gap in how current MLLMs decode creatively encoded meaning through cross-concept relations.

Key Points
  • C4-Eval comprises 184 synthetic and 37 human-created Chengyu items, yielding 884 test cases across 5 task settings.
  • Best closed-source MLLMs reached 50.7% and 48.0% primary accuracy; open-source models scored much lower.
  • Candidate constraints sharply boosted accuracy, while bridge hints and explanation requests offered only modest improvements.

Why It Matters

Creative reasoning is vital for human-AI collaboration; this benchmark reveals MLLMs still struggle with non-literal conceptual understanding.

📬 Get the top 10 AI stories daily