AIES study: Personalized AI needs interaction-level, user-centered auditing
Static audits miss emergent harms that evolve through real user interactions — here's why.
Personalized generative AI systems are adapting to individual users in real time, but the auditing methods meant to catch harmful behavior are still stuck in static, one-size-fits-all evaluations. In a position paper accepted at the Ninth AAAI/ACM Conference on AI, Ethics, and Society (AIES 2026), researcher Hannah Cha argues that these traditional audits systematically miss emergent harms that arise through ongoing interaction with a specific user. Cha identifies three presuppositions baked into current harm auditing paradigms: that harms can be (1) specified outside real-world interaction, (2) defined non-pluralistically within groups, and (3) treated as static. When a model has a personalized history with a user, harm can surface in subtle, context-dependent ways that no simulated test or group-level metric can capture.
The paper also cautions against a tempting fix—letting the system learn each user's personal harm definitions through deeper personalization. That approach, Cha argues, risks shifting the burden of articulating harm onto marginalized users, forcing them to disclose sensitive information and do unpaid safety labor. Instead, she proposes reframing harm as an adaptive, user- and community-centered process, and calls for auditing infrastructures that support ongoing articulation of harm during interaction, not just retrospective evaluation. Design directions include interaction-level audit logs, user-in-the-loop feedback mechanisms, and community-defined safety norms. For teams building personalized chatbots, agents, or recommender systems, the paper is a reminder that safety assurance can't stop at offline benchmarks—it has to live where the model meets the user.
- Argues static, group-level audits cannot detect emergent harms in personalized generative AI systems (arXiv:2608.14692).
- Three flawed presuppositions: harms exist outside real interaction, are uniform within groups, and remain static over time.
- Proposes shifting to interaction-level, user-centered auditing with ongoing articulation instead of retrospective evaluation.
Why It Matters
AI safety teams must rethink audit methods as personalized agents evolve, making user interaction central to harm detection.