C4 benchmark reveals MLLMs fail at creative cross-concept reasoning
Best closed models score only 50.7% on decoding creative Chinese idioms...
A new arXiv paper introduces C4, a cognition-inspired framework for evaluating cross-concept creativity using Chengyu (Chinese idioms). The C4-Eval set includes 184 synthetic items and 37 human-created figures, instantiated across 5 task settings for 884 cases. Testing 10 MLLMs, the strongest closed models reached only 50.7% and 48.0% accuracy, while open-source models lagged substantially behind. Candidate constraints improved accuracy sharply, but bridge hints and explanation requests gave only modest gains — exposing a major gap in how current MLLMs decode creatively encoded meaning through cross-concept relations.
- C4-Eval comprises 184 synthetic and 37 human-created Chengyu items, yielding 884 test cases across 5 task settings.
- Best closed-source MLLMs reached 50.7% and 48.0% primary accuracy; open-source models scored much lower.
- Candidate constraints sharply boosted accuracy, while bridge hints and explanation requests offered only modest improvements.
Why It Matters
Creative reasoning is vital for human-AI collaboration; this benchmark reveals MLLMs still struggle with non-literal conceptual understanding.