Activation steering identifies an internal direction or pattern associated with a behavior and modifies model activations while the model processes a request. The intervention can strengthen, weaken, add, remove, or replace a signal without retraining every parameter.
Steering can test interpretability hypotheses and alter properties such as style, topic, refusal, honesty, or task strategy. Effects may be broad, unstable, or entangled with other behavior, so experiments need controls, dose testing, and evaluation across varied inputs.
ELI5
Activation steering changes a temporary signal inside an AI model to influence what it does. It is like gently adjusting an internal control while the model is working.
For example, researchers can strengthen a signal linked with a cautious response and see whether the model becomes more careful across many prompts. They must also check that unrelated abilities are not damaged.
