Andrew Dai argues that multimodal modelsA multimodal model can process or generate more than one kind of data, such as text, images and audio. can recognize visual patterns while failing to reason reliably about counts, spatial relationships and changes over time. Andrew Dai uses chess, board-game and object-trackingComputer vision enables software to analyze visual inputs such as images or video and extract information about their contents or structure. examples to distinguish recognition from deliberate visual reasoning.
Andrew Dai describes Elorian's approach to multimodal data, synthetic trainingSynthetic data is generated rather than directly observed data, used for purposes such as training or testing AI systems., architecture and visual chains of thought. RoboticsEmbodied AI perceives and acts through a physical or simulated body while interacting with an environment., construction monitoring and CAD illustrate intended applications rather than independently verified production capability, and no safety guarantees are established.
Watch on YouTube




