What is activation steering?

Definition

Activation steering identifies an internal direction or pattern associated with a behavior and modifies model activations while the model processes a request. The intervention can strengthen, weaken, add, remove, or replace a signal without retraining every parameter.

Steering can test interpretability hypotheses and alter properties such as style, topic, refusal, honesty, or task strategy. Effects may be broad, unstable, or entangled with other behavior, so experiments need controls, dose testing, and evaluation across varied inputs.

ELI5

Activation steering changes a temporary signal inside an AI model to influence what it does. It is like gently adjusting an internal control while the model is working.

For example, researchers can strengthen a signal linked with a cautious response and see whether the model becomes more careful across many prompts. They must also check that unrelated abilities are not damaged.

Frequently asked questions

Does activation steering retrain the whole model?

No. It usually changes temporary internal activations during inference rather than permanently updating all learned parameters.

What are the risks of activation steering?

A steering direction can affect unintended behaviors, fail on unfamiliar inputs, reduce capability, or appear interpretable while representing a more complex internal pattern.

Videos explaining activation steering

  1. Tim Scarfe and Tom McGrath beside the words Design AI from the Inside