How Researchers Read AI's Hidden Thoughts

Google DeepMind53:06
0 comments ยท 0 votesOpen discussion
Sign in to join the discussion

    Video summary

    Google DeepMind interpretability lead Neel Nanda describes neural networks as systems that are grown through training rather than explicitly designed. Interpretability tries to reverse engineer the structures and algorithms that emerge, much as biology studies mechanisms produced by evolution.

    Nanda explains why chain-of-thought is useful but incomplete, comparing it to a scratchpad that can expose cheating or confusion without revealing every internal step. He then walks through white-box techniques including probes, which test for chosen concepts, and sparse autoencoders, which search for concepts researchers did not know to request.

    The discussion connects these tools to practical safety work. Cheap probes can help detect harmful use, interpretability can clarify whether models are gaming evaluations, and deeper audits can investigate apparent deception or hidden objectives. Nanda argues for realistic expectations and defense in depth rather than treating interpretability as a complete solution.

    Original YouTube thumbnailWatch on YouTube