Beyond RLHF: From AI Assistance to Automation

AI Engineer18m 5s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Diogo Almeida opens by contrasting striking AI benchmark progress with the continued need for people to supervise many business decisions. He argues that this tension is easier to understand if AI assistance and automation are treated as different goals: interactive tools need to satisfy a human user, while autonomous systems must make dependable decisions without one.

    Almeida traces the assistant paradigm to reinforcement learning from human feedback, or RLHF, which optimizes models using human preferences. In his account, that objective can reward confident, agreeable responses even when a task requires calibrated correctness. He places current coding assistants in the same broad assistance era, while acknowledging their usefulness. These are his interpretations of model behavior, not a universal finding that all RLHF systems fail at automation.

    He proposes a future stack designed for reliable automation and more capable software, and says TypeSafe AI is pursuing a different optimization target. The presentation does not disclose or independently validate that system. During questions, he distinguishes pretraining from post-training and says the intended objective is calibrated decision-making rather than either human preference or simply verifiable answers.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Diogo Almeida in a blue top on black beside the headline ‘BEYOND RLHF’ and blue branching decision lines. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 31 July 2026 and duration 18m 5s.

    Diogo Almeida argues that RLHF shaped highly useful assistants but that reliable automation will require different objectives, calibrated decisions and software designed around autonomous work.