Hidden AI Reasoning Can Be Extracted

Wes Roth35:48
0 comments · 0 votesOpen discussion
Sign in to join the discussion

    Video summary

    Wes Roth begins with research on extracting raw reasoning traces from proprietary language-model APIs. These traces can expose sensitive data that models intentionally keep out of final answers, and they may reveal enough behavioral structure to make distillation attacks against closed models easier. The researchers disclosed the method to affected labs before publishing it.

    The video also covers an Anthropic research model making progress on a lower bound related to the Riemann hypothesis. The system generated millions of tokens, coordinated dozens of subagents, ran thousands of shell commands and tested hundreds of failed ideas before finding an improvement. Roth uses the result to argue that useful human oversight can increasingly mean choosing a problem and encouraging an agentic research process rather than supplying every technical step.

    A second theme is the spread of persistent autonomous agents. Grokbot is described as an agent with its own virtual machine, messaging and delegation, while Meta's Muse models and other open releases show how capable models can reach developers outside the largest closed platforms. Roth contrasts that accessibility with proposals for invisible text watermarking, whose effects on attribution, copyright and model competition remain uncertain.

    The final section examines Meta's argument for broadly available superintelligence and questions whether conventional entrepreneurship still makes sense if systems surpass people across nearly every cognitive task. It also surveys reported leadership changes at Google, delayed model releases and departures connected to military applications, while noting that Google's infrastructure and resources still make it a major AI competitor. The investment sponsorship is omitted.

    Original YouTube thumbnailWatch on YouTube