Kenny Workman describes how large biological experiments create demanding data-analysis workflows and argues that executable analysis can provide a foundation for evaluating scientific agents. He distinguishes writing code or recalling biology facts from extracting a defensible scientific insight from real experimental data.
Kenny Workman describes SpatialBench and its focus on concrete steps in spatial-biology analysis. Tasks combine experimental data, a scientific goal and a deterministic grader. He argues that good tests must be verifiable, remain valid across reasonable analysis paths and require interaction with the supplied data rather than a memorized answer.
Kenny Workman explains why scientists reviewing one another's work became important for checking the benchmark. Ambiguous task wording, discretionary analysis choices and arbitrary numerical thresholds can cause a grader to reject a valid result, so human verification helps expose weaknesses in the evaluation itself.
Kenny Workman discusses longer tasks spanning entire scientific workflows, where a final pass-or-fail reward can be too sparse to reveal useful progress. He describes intermediate rubrics tied to important analysis steps, while acknowledging that their correlation with verified outcomes is not yet strong enough to inspire full confidence for benchmarking or reinforcement learning.
Kenny Workman outlines extensions into other biological data types and drug-discovery evaluations. He also discusses safety tests that distinguish routine scientific requests from misuse, presenting this as an evaluation problem that needs more nuance than broadly refusing biology questions.
Watch on YouTube




