ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2025
Year
… In this work, we propose SPARE, a training-free representation engineering method that uses pre-trained sparse auto-encoders (SAEs) to control the knowledge selection behaviour of …
This paper addresses a critical challenge in large language models: controlling what knowledge the model uses when generating responses. Traditional methods like fine-tuning require significant computational resources and can lead to catastrophic forgetting. SPARE introduces a training-free approach that uses sparse auto-encoders (SAEs) to directly manipulate internal representations, offering a more efficient and interpretable alternative.
The significance lies in its potential to democratize model steering. By leveraging pre-trained SAEs, practitioners can adjust model behavior without access to large-scale compute or original training data. This is particularly relevant for domain-specific applications where fine-tuning is impractical. Moreover, the method aligns with the growing interest in mechanistic interpretability, providing a bridge between understanding and controlling LLMs.
The abstract does not provide specific numerical metrics, but the paper claims that SPARE effectively controls knowledge selection behavior. It compares favorably to fine-tuning baselines, suggesting that it can achieve similar or better performance without the computational overhead. The method is evaluated on benchmarks that likely test factual recall and domain-specific knowledge, though exact numbers are not available in the abstract.
SPARE contributes to the growing field of representation engineering, offering a practical tool for AI practitioners to customize LLM behavior. Its training-free nature makes it accessible for rapid prototyping and deployment. The use of SAEs also enhances interpretability, as the edited features can be inspected to understand what knowledge is being selected. This could lead to more transparent AI systems and facilitate safer deployment in sensitive applications. Future work may extend SPARE to other behavioral aspects beyond knowledge selection, such as style or safety.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba