How Interpretability Can Shape AI from Within

Machine Learning Street Talk1:40:15
0 comments · 0 votesOpen discussion
Sign in to join the discussion

    Video summary

    Tom McGrath describes Goodfire's effort to turn mechanistic interpretability into practical infrastructure for understanding model internals. Instead of treating a neural network as an opaque collection of weights, the work looks for meaningful features, directions and geometric relationships inside its activations.

    McGrath explains that these representations can support controlled generalization and data debugging. Teams can identify which concepts drive a behavior, trace unexpected outputs back to training examples and test interventions that change a model's internal state without relying only on prompts.

    The same tools may help evaluate hallucinations, reward hacking and other safety failures. McGrath stresses that useful interpretability must connect internal measurements with real behavior, while remaining careful about what a discovered feature actually represents.

    Original YouTube thumbnailWatch on YouTube