
Modular Pretraining: A New Approach to Containing Dangerous AI Knowledge
Researchers at Anthropic and AE Studio have introduced Gradient Routed Auxiliary Modules (GRAM), a method that isolates dangerous knowledge in large language models into switchable modules during training. This approach allows operators to control access to sensitive content, potentially reducing risks of misuse. Preliminary experiments show promise across models up to 5B parameters, but the method has not yet been applied to production-scale systems.