New paper defines good explanations and why LLMs struggle with them
Researchers propose a counterfactual definition of explanations that highlights LLMs' unique challenges.
Deep Dive
Mahon, Ford, and Hackett propose a definition of good explanations based on counterfactuals that account for the listener's prior beliefs. They explore the ramifications of this definition for AI explainability and, in particular, why LLM outputs are difficult to produce good explanations for.
Key Points
- Defines good explanations as counterfactual and dependent on the listener's prior beliefs
- LLMs produce outputs without causal reasoning, making true explanations difficult
- Framework suggests explainability methods must model user knowledge and model limitations
Why It Matters
Exposes a core transparency gap in LLMs, pushing researchers toward more rigorous explainability standards.