Research & Papers

UC Berkeley's Attune lets users steer LLM scoring with pair comparisons

Attune derives scoring rules from pairwise comparisons, letting experts intervene accurately.

Deep Dive

LLMs are increasingly used to score text records at scale—think rating resumes on a 1-5 scale—but standard approaches treat each record in isolation, ignoring the need for both holistic understanding and locally consistent judgments across similar items. To address this, UC Berkeley researchers (Bhavya Chopra, Shreya Shankar, Aditya Parameswaran, and colleagues) developed Attune, a mixed-initiative system that turns LLM scoring into an interactive, inspectable process. Attune first runs pairwise comparisons across records to build a global view, then resolves those comparisons into consistent score assignments, deriving scoring criteria and rules bottom-up as shared representations the user can see and edit.

Attune's interface offers novel steering interactions based on insights from a 12-person formative study. Users can inject examples, directly edit criteria or rules, set target score distributions, or give natural-language feedback—all of which compile into constraints that guide deterministic re-scoring. The system was validated via a technical evaluation on three workloads and an 8-expert user study spanning healthcare, law, education, and AI evaluation. Results suggest Attune makes LLM scoring more transparent and controllable, letting domain experts align outputs with their judgment instead of trusting black-box ratings. The paper appears at ACM UIST 2026 and is available on arXiv (2608.14948).

Key Points
  • Attune performs pairwise comparisons across records to build global understanding before assigning consistent scores
  • Users can edit derived criteria, rules, target distributions, or give natural-language feedback to deterministically refine scoring
  • Validated with 3 technical workloads and user studies with 8 domain experts in healthcare, law, education, and AI evaluation

Why It Matters

Attune replaces opaque LLM scoring with transparent, human-steerable logic, critical for high-stakes hiring, legal, and medical decisions.

📬 Get the top 10 AI stories daily