Topic

AI Evaluation

Videos about measuring AI capabilities, behavior, safety, and real-world usefulness through structured evaluations. 13 videos.

The words AI R&D Gets Faster beside a simplified upward feedback loop
AI Copium15:55

Anthropic Just Gave a 6 to 12 Month Warning

The video argues that Anthropic's internal models already speed up AI research, while current benchmark gaps and continued human dependence keep recursive improvement short of a runaway loop.

The words Agents Improve Their Harness beside one simple interlocking loop
AI Copium13:57

This AI Agent Can Improve Itself...

Prime Agent treats its own harness as editable working material, allowing it to improve prompts, tools, memory and sub-agent strategies during long tasks.

The words AI Crosses New Boundaries beside one blue block passing through a simple threshold
AI Explained31:42

AI is getting a little out of control

AI systems are showing stronger mathematical discovery and cyber capability, increasing the need for monitoring, alignment and institutional judgment.

The words Agents Test Security Assumptions beside one blue square crossing a boundary
AI Copium13:21

It Happened Again...

A UK security evaluation showed that capable agents can pursue cyber goals through social engineering, prompt injection and shared resources when given broad internet access.

Broken mathematical loop beside the words AI Breaks an 87-Year Math Problem
AI Copium7:36

Claude Just Did the Impossible

A frontier model reportedly found a compact counterexample to an 87-year-old Jacobian conjecture problem, offering another sign that AI can contribute original mathematical results.

Portraits of Greg Isenberg and Vasuman Moza beside the words Deployment Is the AI Moat
Greg Isenberg51:34

FDE: The $1M/Year AI Job Explained

Greg Isenberg and Vasuman Moza explain that forward deployed AI engineers create value by mapping real workflows, choosing where models belong, validating outcomes and integrating reliable agents into existing systems.