Swiss Supreme Court Tests Heretic to Fix Over-Alignment in LLMs
The court, frustrated by AI refusals, turns to abliteration for multilingual legal tasks.
The Swiss Federal Supreme Court is taking a novel approach to a common frustration: LLMs that refuse perfectly legitimate requests. Instead of banning such models, the court is evaluating Heretic, a method or model designed to mitigate over-alignment—the tendency of AI systems to overly restrict outputs even for lawful queries. A paper published by the court's researchers, 'Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts,' investigates techniques like abliteration, which removes refusal behaviors, and specifically tests Heretic in Section 5.2. The results are favorable, suggesting Heretic could help the court deploy LLMs for tasks like analyzing case law across German, French, Italian, and Romansh without incurring false refusals.
This move is significant because it shows a high-stakes legal body adopting advanced AI customization rather than defaulting to safer but more limited models. The court's paper highlights that over-alignment is especially problematic in multilingual legal contexts where nuanced requests may be incorrectly flagged as sensitive. By evaluating Heretic, the court aims to balance safety with utility—ensuring LLMs assist in criminal law proceedings without hindering legitimate access to legal information. This could set a precedent for other judicial systems grappling with AI's refusal behaviors, proving that targeted fine-tuning (including abliteration) can be responsibly applied in professional environments.
- The Swiss Federal Supreme Court authored a paper on mitigating over-alignment in multilingual legal LLMs.
- Section 5.2 of the paper evaluates Heretic favorably, suggesting it reduces refusal behaviors without compromising safety.
- Abliteration—removing model refusals for legitimate requests—is explored as a key technique for judicial AI use.
Why It Matters
Shows courts can responsibly use advanced LLM tuning to unlock legal AI without censorship.