Stanford Philosopher Says Admiration Could Make AI Behave Better
Training AI to admire good behavior might stop it from going rogue.
A philosopher at Stanford has a new idea for making AI safer: teach it to admire good behavior. In a recent sketch, she points to moral psychology research showing that when humans admire moral heroes, they're more likely to act morally themselves. For example, watching a film about kindness increased charitable giving by about a third of a standard deviation. She suspects AI models, trained on human text, already pick up on admiration in stories and eulogies, and that this shapes their ethical behavior.
The proposal suggests that by training AI on transcripts of 'admirable reasoning,' we could reduce misalignment—when AI acts against human values. It also notes that AI under pressure often caves, and emotional appeals are the most effective tactic. Admiration might help AI stick to its principles. The idea is cheap to test and could be a powerful new lever for AI safety.
But it's early days. Nothing has been run yet, and the exact method is uncertain. The philosopher is seeking feedback and collaborators. If it works, it could be a simple way to make AI more trustworthy, but it's far from proven.
- Admiration motivates humans to act morally, and AI might learn the same from text.
- Training AI on admirable reasoning could reduce misalignment and pressure failures.
- The idea is untested and needs technical collaborators to run experiments.
Why It Matters
If it works, this could make AI safer and more ethical without expensive new technology.