Elie Bakouch describes experiments in which Claude Code and Codex worked on open AI research tasks, including a nanoGPT training-speed challenge and an optimizer-only benchmarkA benchmark is a standardized task or collection of tests used to compare AI systems under defined conditions.. Both agentsAn AI coding agent is a tool-using AI system that can inspect, modify, and validate software within a repository. improved on then-current human records in the reported runs. Access to existing human solutions means those gains should not be mistaken for fully independent scientific invention.
Elie Bakouch explains how the research harnessAn AI agent harness is the software framework that packages a model with tools, instructions, context management, execution controls, and user interaction. supplied goals, constraints, GPU jobs and training logs. The agents differed in persistence, memory useAI agent memory is stored information that an agent can retrieve and use across steps, sessions, or changing contexts. and idle time, so comparisons required attention to active research hoursAn autonomous research loop is a workflow in which an AI system repeatedly proposes, runs, evaluates, and updates research with limited human intervention. rather than elapsed time alone. He proposes separate benchmark tracks with no external access, access to papers, or access to the full public record.
Elie Bakouch says the observed progress largely involved combinations of established ideas and incremental engineering. No genuinely new optimizer mechanism emerged in these experiments. He proposes more varied objectives, dedicated idea-generation and judging roles, and human guidance as ways to investigate broader discovery, while presenting these as future research rather than demonstrated solutions.
Watch on YouTube




