What is interpretability?

Definition

Interpretability methods investigate the internal representations, computations, or decision factors associated with an AI system's output. Some methods explain a particular result, while others look for broader mechanisms or recurring patterns inside a model.

Interpretability can help researchers diagnose failures, test safety hypotheses, and form better questions for evaluation. An explanation is still evidence with limits rather than complete access to a model's reasoning, so it should be validated and combined with behavioral testing.

ELI5

Interpretability is an attempt to understand what is happening inside a complicated AI instead of judging it only by the final answer. It is similar to examining both a machine's output and the parts that influenced it.

For example, a researcher might test which internal features change when an AI recognizes a risky instruction. That clue can help investigation, but it does not automatically explain every decision.

Is interpretability the same as a model explaining itself?

Not necessarily. A generated explanation may be inaccurate, while interpretability research uses methods designed to study behavior or internal mechanisms directly.

Can interpretability guarantee safe behavior?

No. It can provide useful evidence and reveal problems, but it does not replace testing, monitoring, or other safeguards.

  1. Theo Browne and Dario Amodei beside the headline Slow the AI Frontier? on a black background.
    I think they mean it this time
    Theo34m 48s3 VIEWS
  2. Nathaniel Whittemore with Dario Amodei and Sam Altman against a black background framing the debate over pacing frontier AI.
  3. Theo Browne in a blue top beside the headline FRONTIER PACING on a black background.
Definition card for interpretability: What is interpretability?

Interpretability is the study of methods that help people understand how an AI system represents information and produces its behavior.