Reilly Haskins describes METR's live per-action monitorAn AI monitoring agent observes systems or information sources, interprets signals, and reports or responds to relevant changes. in a discussion with Theo Jaffee. Before an evaluation agent executes a proposed action, another language model assesses it for possible real-world harm. Suspicious actions halt for human review, allowing a researcher to approve the action or stop the run.
Reilly Haskins argues that a monitor is only useful if it covers the relevant inference. Local models, external APIs and inconsistent logging can leave gaps. Structured transcripts and explicit role boundariesAn agent role boundary defines the specific responsibility, tools and decisions assigned to an AI agent. help resist prompt injection, but no known formatting choice eliminates jailbreaks or deceptive evasion. Red-teamingAdversarial evaluation deliberately searches for inputs and edge cases that reveal where an AI system's rules or behavior fail. and documenting a control system's assumptions remain essential.
The discussion distinguishes increasingly opaque model reasoning from action-level oversight. Reilly Haskins notes that stronger models may reason without a readable chain of thought, but monitor models may also improve. Cybersecurity tasks, very difficult tasks and more capable agents raise evaluation risks; automated enforcement through a central inference platform can reduce reliance on researchers interpreting written policies correctly.
Reilly Haskins discusses alert fatigue and the value of reviewers who understand a task's normal behavior. The current implementation's reported cost and latency overheads might be reduced with cheaper first-stage monitors, escalation hierarchies or less frequent checks, though these choices involve safety tradeoffs. He recommends decomposing monitoring claims into explicit subclaims and examining the evidence behind each one.
Watch on YouTube




