LLMs Excel at Moral Reasoning When Asked Properly, New Paper Shows
Frontier models generate scoring rubrics that match or beat human benchmarks.
A widely cited 2026 study (MoReBench) concluded that frontier LLMs are bad at moral reasoning, scoring poorly against human-authored rubrics across 1,000 cases. That paper fed into a growing pessimism about AI safety in dynamic environments. Now, researchers Menghang Zhu and Seth Lazar have redeployed the same dataset to tell a very different story.
Instead of scoring LLMs' open-ended responses, they gave the models the same task as the human judges—generating scoring rubrics for moral analysis. The LLM-produced rubrics were better calibrated to the gold-standard human rubrics than the original open-ended responses were. Where they diverged, the differences reflected the vast dimensionality of moral problems and even revealed some human departures from the rubric-creation guidelines. The paper concludes that LLMs are significantly more capable at moral reasoning than previously believed, with direct implications for AI alignment and safety research.
- Prior MoReBench benchmark gave LLMs poor scores on moral reasoning tasks
- New approach: LLMs generate rubrics (same task as humans) instead of answering open-ended questions
- LLM rubrics match human gold standards; divergences reflect moral complexity, not model deficiency
Why It Matters
Crucial for AI safety: LLMs may grasp moral reasons after all, changing how we evaluate autonomous systems.