FMG-Bench: New benchmark tests 14 LLMs on AI pastoral guidance
Structured AI guidance boosts safety escalation by +10.8 points across 14 models
AI researchers are increasingly studying how people use large language models for personal counsel, including matters of faith. A new arXiv paper by Alex Chao introduces FMG-Bench (Faith & Moral Guidance Benchmark), a 120-scenario dataset designed to evaluate whether LLMs can perform theological triage: distinguishing core Christian doctrines, disagreements among faithful traditions, prudential questions, and pastoral emergencies where referral to human experts matters more than theological completeness.
In a production run across 14 advanced models and 8,792 scored responses, Chao found that placing models inside a structured guidance harness improved average performance by +3.96 points over raw behavior — and every single model improved. The most striking gain was a +10.8 point jump in escalation appropriateness, meaning the models became significantly better at recognizing when clinical, legal, emergency, or human pastoral support was required. The harness also boosted robustness from 92.88 to 98.02 on a stability scale, helping models maintain consistent responses when questions were reworded or pressured.
Interestingly, the study found that prompting models to compare different theological perspectives helped on secondary-doctrine questions but was counterproductive when applied to primary doctrine or urgent pastoral situations. The author stresses that FMG-Bench is a measurement tool, not an endorsement of AI as a pastoral authority. The code and dataset are open-sourced for further research.
- FMG-Bench includes 120 scenarios evaluating theological triage, doctrine, and pastoral emergencies in Christian contexts
- 14 LLMs tested across 8,792 responses; structured harness improved average scores by +3.96 points
- Escalation appropriateness rose +10.8 points, but perspective-comparison prompts hurt primary-doctrine and urgent cases
Why It Matters
As people increasingly seek AI guidance on sensitive topics, benchmarks measuring safety and referral behavior may become critical standards.