Which character are we evaluating? Persona stability and AI welfare
Which character are we evaluating? Persona stability and AI welfare
Deep Dive
TL;DR: AI welfare is hard to evaluate when one model can inhabit many personas. Recent work suggests that future training may produce a single stable underlying persona that can play many roles, making model welfare much easier to evaluate. One of the hardest questions in AI welfare is deciding what