Research & Papers

New 'Weights to Words' method lets users inspect and edit AI preference models in natural language

This technique makes black-box preference models transparent and editable in real-time.

Deep Dive

Zachary Wojtowicz, Ayush Nayak, and Jacob Andreas introduce "weights to words," a method that automatically discovers domain-relevant preference dimensions from choice data, each described in natural language and paired with a vector in the model's representational space. This allows users to inspect and edit inferences in real time. The method is illustrated on four domains: moral dilemmas, movies, wines, and free-form LLM responses. Two pre-registered experiments (N=450 on moral dilemmas, N=449 on movie selection) show that regularizing a preference model toward the learned basis improves prediction accuracy on held-out choices, and incorporating participants' structured edits further improves accuracy. Participants preferred the method's inferred preference profiles and endorsed its predictions as more accurate.

Key Points
  • Automatically discovers 5-10 preference dimensions per domain, each described in natural language (e.g., 'utilitarian vs deontic' for moral dilemmas).
  • Two pre-registered experiments with 450+ participants each show prediction accuracy gains of 3-5% from regularization and 2-4% additional gains from user edits.
  • Demonstrated on four domains: moral dilemmas, movies, wines, and free-form LLM responses; paper is 42 pages with 22 figures and 14 tables.

Why It Matters

Makes AI preference models transparent and correctable by humans, improving trust and alignment in high-stakes decisions.

📬 Get the top 10 AI stories daily