AI Safety

When talked into harm, a model blames the answer, not itself (an interpretability study of guilt vs shame)

When talked into harm, a model blames the answer, not itself (an interpretability study of guilt vs

Deep Dive

x This website requires javascript to properly function. Consider activating javascript to get access to all site functionality. When talked into harm, a model blames the answer, not itself (an interpretability study of guilt vs shame) — LessWrong Activation Engineering Emergent Misalignment Interpr

📬 Get the top 10 AI stories daily