Anthropic's GRAM method controls dangerous AI knowledge with modular switches
New GRAM technique lets AI models toggle dangerous knowledge on/off like light switches
Deep Dive
AE Studio, in collaboration with Anthropic, introduces GRAM (Gradient Routed Auxiliary Modules), a method to isolate dangerous knowledge in language models into switchable modules. Tested on models from 50M to 5B parameters, a single GRAM model can be reconfigured to match the performance of any of five distinct filtered models. It enables access
Key Points
- GRAM isolates dangerous knowledge into switchable auxiliary modules, tested on models up to 5B parameters.
- A single GRAM model can approximate the performance of five separately trained data-filtered models.
- On an 800M model, GRAM successfully controlled access to virology, cybersecurity, nuclear physics, and specialized code domains.
Why It Matters
Enables dynamic, fine-grained access control to AI capabilities without retraining, reducing misuse risk while preserving performance.