James Shi presents DeepSWE as a benchmark intended to test unfamiliar software-engineering work without relying on public pull-request solutions. He describes 113 newly authored tasks across 91 repositories and five programming languages. Removing publicly available answers reduces one source of contamination, but does not make any benchmark a complete measure of engineering ability.
Shi examines how evaluation results can hide weaknesses. Closely related tasks can overweight a narrow capability, while tests tied to helper names or a specific implementation can reject a valid solution. He contrasts exploratory and literal agent behavior and discusses cases where agents miss multipart requirements or find answers in retained Git history. The reported leaderboard reflects the presentation's evaluation setup, not a current universal ranking.
The talk emphasizes tasks written by contributors who understand their repositories. Concise, high-level prompts require agents to explore existing conventions and coordinate larger changes without being handed a prescribed implementation. Shi argues that verifiers should test requested observable behavior rather than a particular private helper structure.
Shi also distinguishes underlying model capability from the prompts, tools and runtime surrounding it. A brief instruction suggesting that tests are already handled can discourage self-verification. A common harness enables a more controlled comparison, but differences from native agent products and the task mix remain limitations.
For the benchmark's next iteration, Shi describes separating agent and evaluator environments, removing unnecessary Git references and improving standardized reports. Debugging, refactoring, broader repositories and less prescriptive evaluation remain areas for expansion. These changes are methodological proposals and reported work, not proof that benchmark scores guarantee production reliability.
Watch on YouTube




