David Brumley, a Carnegie Mellon University professor and Bugcrowd's Chief AI and Science Officer, explains how his team designs reinforcement-learning environments for cybersecurity. He frames training along two axes: harder target programs and harder levels of exploitation. A student-learning story illustrates the progression from easier exercises to demanding real-world targets, but the technical focus is on how to measure what a model actually learned.
The proposed environments place vulnerable software in reproducible containers and use deterministic graders rather than letting an AI judge its own success. David Brumley argues that a crash is a useful early signal, but not proof that a model could control the program. He also explains why benchmarks built around a single known bug can reward a model for finding the easiest flaw repeatedly or can leak the answer by pointing to the vulnerable function.
His team's audit-task design instead asks for distinct proofs across all vulnerabilities a model can find, including ones the benchmark authors did not know about. A grader separates duplicate crashes and uses precision and recall to discourage both repetition and unsupported claims. David Brumley presents this as a way to use real open-source software while keeping the reward signal tied to demonstrated behavior.
For a harder evaluation, David Brumley reports results from 41 previously verified V8 vulnerabilities. The talk distinguishes triggering a crash from reaching stronger exploitation milestones and says the latter separates model performance much more sharply. These are speaker-reported findings; the official talk notes discrepancies in the displayed chart, so public copy does not assert exact model percentages without parent verification.
The conclusion raises an open-science and safety tension: publishing a benchmark is useful, but releasing complete transcripts of newly produced high-value exploits could create risk. David Brumley says his team is building further training environments from previously unseen vulnerabilities to reduce memorization. The summary omits exploit mechanics and direct calls to download or run the benchmark.
Watch on YouTube




