Research & Papers

Gemma 2 2B Study: Active SAE Planes Show Less Holonomy, Reversing Prediction

Preregistered test on Gemma 2 2B finds active feature planes have less geometric curvature.

Deep Dive

A new preregistered study by Larry Richards on Gemma 2 2B tests whether geometric curvature (holonomy) concentrates on active sparse-autoencoder (SAE) feature planes. The paper operationalizes the broader 'semantic-concentration prediction' by measuring holonomy at the residual-stream readout from layer 12 to 13. Using a restricted-Jacobian transport rule, the author carries local frames around small loops and normalizes rotation by area. The design, thresholds, analysis, and verdict rules were all frozen before measurements were inspected.

The result was a clear reversal: active-feature planes carried significantly less holonomy than mixed-feature controls, with an adjusted log contrast of -0.29439 (95% CI [-0.43989, -0.14889]). Post-freeze diagnostics supported the area law on a small validation subset and identified transport distortion as a potential confound. The author emphasizes this is a narrow, auditable operational reversal—not a causal claim that meaning suppresses holonomy. Competing explanations include activation-strength geometry, degree of feature engagement, dictionary geometry, and transport shear.

Key Points
  • Preregistered confirmatory study on Gemma 2 2B, with all design and analysis rules frozen before data inspection
  • Active SAE feature planes carried less holonomy than mixed-feature controls (log contrast -0.29439, 95% CI [-0.43989, -0.14889])
  • Cause remains open; multiple alternatives include activation-strength geometry, dictionary geometry, and transport shear

Why It Matters

Challenges assumptions about geometric structure in SAEs, underscoring the need for causal interpretability research.

📬 Get the top 10 AI stories daily