Steering knowledge selection behaviours in LLMs via sae-based representation engineering
Unknown
SPARE is a training-free method using pre-trained sparse auto-encoders to steer knowledge selection in LLMs by editing internal representations.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
SPARE is a training-free method using pre-trained sparse auto-encoders to steer knowledge selection in LLMs by editing internal representations.
Unknown
SHARP is a representation-level intervention framework that modulates hallucination in LVLMs by steering internal representations.
Unknown
This paper introduces In-context Vectors (ICV) to make in-context learning more effective and controllable by steering latent space, offering a computationally efficient alternative.
Eric J. Bigelow, Daniel Wurgaft, YingQiao Wang, et al.
A unified Bayesian model shows that activation steering and in-context learning control LLMs by altering concept priors and accumulating evidence, respectively.