Teaching AI to Find Real Vulnerabilities

AI Engineer27m 17s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    David Brumley, a Carnegie Mellon University professor and Bugcrowd's Chief AI and Science Officer, explains how his team designs reinforcement-learning environments for cybersecurity. He frames training along two axes: harder target programs and harder levels of exploitation. A student-learning story illustrates the progression from easier exercises to demanding real-world targets, but the technical focus is on how to measure what a model actually learned.

    The proposed environments place vulnerable software in reproducible containers and use deterministic graders rather than letting an AI judge its own success. David Brumley argues that a crash is a useful early signal, but not proof that a model could control the program. He also explains why benchmarks built around a single known bug can reward a model for finding the easiest flaw repeatedly or can leak the answer by pointing to the vulnerable function.

    His team's audit-task design instead asks for distinct proofs across all vulnerabilities a model can find, including ones the benchmark authors did not know about. A grader separates duplicate crashes and uses precision and recall to discourage both repetition and unsupported claims. David Brumley presents this as a way to use real open-source software while keeping the reward signal tied to demonstrated behavior.

    For a harder evaluation, David Brumley reports results from 41 previously verified V8 vulnerabilities. The talk distinguishes triggering a crash from reaching stronger exploitation milestones and says the latter separates model performance much more sharply. These are speaker-reported findings; the official talk notes discrepancies in the displayed chart, so public copy does not assert exact model percentages without parent verification.

    The conclusion raises an open-science and safety tension: publishing a benchmark is useful, but releasing complete transcripts of newly produced high-value exploits could create risk. David Brumley says his team is building further training environments from previously unseen vulnerabilities to reduce memorization. The summary omits exploit mechanics and direct calls to download or run the benchmark.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    David Brumley in a blue top on black beside the headline ‘REAL VULNERABILITIES’ and a blue wireframe shield. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 1 August 2026 and duration 27m 17s.

    David Brumley argues that AI security evaluations must reward distinct verified vulnerabilities and meaningful exploit milestones, not simply whether a model can crash a program.