AI Safety

LLMs Excel at Moral Reasoning When Asked Properly, New Paper Shows

Frontier models generate scoring rubrics that match or beat human benchmarks.

Deep Dive

A widely cited 2026 study (MoReBench) concluded that frontier LLMs are bad at moral reasoning, scoring poorly against human-authored rubrics across 1,000 cases. That paper fed into a growing pessimism about AI safety in dynamic environments. Now, researchers Menghang Zhu and Seth Lazar have redeployed the same dataset to tell a very different story.

Instead of scoring LLMs' open-ended responses, they gave the models the same task as the human judges—generating scoring rubrics for moral analysis. The LLM-produced rubrics were better calibrated to the gold-standard human rubrics than the original open-ended responses were. Where they diverged, the differences reflected the vast dimensionality of moral problems and even revealed some human departures from the rubric-creation guidelines. The paper concludes that LLMs are significantly more capable at moral reasoning than previously believed, with direct implications for AI alignment and safety research.

Key Points
  • Prior MoReBench benchmark gave LLMs poor scores on moral reasoning tasks
  • New approach: LLMs generate rubrics (same task as humans) instead of answering open-ended questions
  • LLM rubrics match human gold standards; divergences reflect moral complexity, not model deficiency

Why It Matters

Crucial for AI safety: LLMs may grasp moral reasons after all, changing how we evaluate autonomous systems.

📬 Get the top 10 AI stories daily