The Best Models Still Reason Like Toddlers - Andrew Dai, Elorian

AI Engineer18m 16s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Andrew Dai argues that multimodal modelsA multimodal model can process or generate more than one kind of data, such as text, images and audio. can recognize visual patterns while failing to reason reliably about counts, spatial relationships and changes over time. Andrew Dai uses chess, board-game and object-trackingComputer vision enables software to analyze visual inputs such as images or video and extract information about their contents or structure. examples to distinguish recognition from deliberate visual reasoning.

    Andrew Dai describes Elorian's approach to multimodal data, synthetic trainingSynthetic data is generated rather than directly observed data, used for purposes such as training or testing AI systems., architecture and visual chains of thought. RoboticsEmbodied AI perceives and acts through a physical or simulated body while interacting with an environment., construction monitoring and CAD illustrate intended applications rather than independently verified production capability, and no safety guarantees are established.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Andrew Dai beside the headline AI STILL CAN'T SEE on a black background. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 22 September 2026 and duration 18m 16s.

    Andrew Dai distinguishes visual recognition from reasoning and presents Elorian's approach to more grounded multimodal models.