When talked into harm, a model blames the answer, not itself (an interpretability study of guilt vs shame)
When talked into harm, a model blames the answer, not itself (an interpretability study of guilt vs
Deep Dive
x This website requires javascript to properly function. Consider activating javascript to get access to all site functionality. When talked into harm, a model blames the answer, not itself (an interpretability study of guilt vs shame) — LessWrong Activation Engineering Emergent Misalignment Interpr