AI Safety

Which character are we evaluating? Persona stability and AI welfare

Which character are we evaluating? Persona stability and AI welfare

Deep Dive

TL;DR: AI welfare is hard to evaluate when one model can inhabit many personas. Recent work suggests that future training may produce a single stable underlying persona that can play many roles, making model welfare much easier to evaluate. One of the hardest questions in AI welfare is deciding what

📬 Get the top 10 AI stories daily