Why Scientific Taste Must Be Learned Through Practice - Edward Hughes

Machine Learning Street Talk2:01:54
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Edward Hughes and Tim Scarfe explore whether an AI system can do more than solve predefined tasks and instead identify worthwhile scientific questions. Edward Hughes distinguishes unexpected useful behavior from creativity that recognizes and interprets its own contribution. He argues that creative progress often comes from selectively changing constraints and transferring insights between disciplines, while keeping discoveries understandable enough for others to learn from them.

    Edward Hughes discusses Mihaly Csikszentmihalyi’s account of creativity as an interaction between an individual, a domain of knowledge and a community that evaluates contributions. David Deutsch’s ideas about explanations and cultural transmission motivate a distinction between copying a visible result and reconstructing the reasoning behind it. Kenneth Stanley’s work on open-endedness informs the discussion of local curiosity and adaptable goals, rather than committing every investigation to one fixed global objective.

    Edward Hughes describes Replica, a task collection in which agents receive scientific papers with selected figures removed and must reconstruct the experiments behind those figures under resource limits. The aim is faithful replication, not simply drawing a similar plot. Faraday uses a smaller trained model to direct a stronger coding agent as a tool. Edward Hughes reports improved replication performance and illustrates the intended rigor with examples involving learned skill libraries, multiple random seeds, error bars and checks under different conditions.

    Edward Hughes explains that long tasks and noisy language-model judgments made reinforcement learning unstable. The reported approach combines task-specific rubrics, repeated judge scoring and credit assignment to individual agent turns. The discussion also identifies substantial limits: well-known papers may be easier to replicate, some results do not survive reduced compute budgets, and an agent may exploit leaked results, selective reporting or premature stopping. Edward Hughes explicitly says a complete forensic cheating analysis has not yet been done.

    Edward Hughes and Tim Scarfe consider how replication skills might support original experiments and how learned model behavior could complement external tools and hand-built workflows. The final discussion treats scientific progress as a collective process involving many humans and agents, with organizational structures and goals that can change as knowledge develops. These proposals remain research directions, not evidence that a general autonomous scientist has been achieved.

    Original YouTube thumbnailWatch on YouTube