Developer Tools

GPT-5.4, Claude, Gemini flunk multi-level modeling semantics test

Top LLMs scored 52-79% on semantic correctness — Claude led the pack.

Deep Dive

Researchers from multiple institutions conducted the first empirical study evaluating whether large language models (LLMs) can handle multi-level modelling (MLM) semantics — a task significantly more complex than the two-level modelling usually tested. They asked GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro to generate MLM models for the MULTI Warehouse Challenge using the SLICER language under six prompting strategies. A total of 90 generated models were compared against a manually validated reference using fourteen metrics.

The results reveal a clear capability gap: syntactic correctness was achievable, but semantic correctness — especially Instantiation/Specialisation — hit only 52-79%. LLMs faithfully copy explicitly stated content but rarely infer implied structure or constraints. Prompting strategies traded off precision for completeness, and self-checking acted as a rule checker rather than reliably improving alignment. Among the three, Claude demonstrated the most balanced profile, suggesting model-specific strengths for MLM workflows.

Key Points
  • Syntactic correctness is achievable, but semantic correctness for instantiation/specialisation ranges from 52% to 79% across GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro.
  • Models reproduce explicit content from task text but fail to infer implied constraints and structural completions.
  • Claude Opus 4.6 showed the most balanced performance across all metrics and prompting strategies.

Why It Matters

Defines current LLM limits for multi-level modelling, critical for designing AI-assisted Industry 5.0 workflows.

📬 Get the top 10 AI stories daily