Bilawal Sidhu explains how SAM 3 expands segmentation from fixed object categories to open-ended concepts described in natural language. It can find every instance of an object, state, or action across images and video, including queries such as people sitting or a person holding a package, then track those instances over time.
SAM 3D adds geometry. Its object model reconstructs scene elements as 3D assets, while its body model estimates human shape and pose. Used together, the models connect pixels, semantic concepts, camera pose, and spatial relationships, providing the perception needed for manipulation, navigation, augmented reality, and automated visual-effects workflows.
The central scaling idea is a data engine that asks humans to rank model-generated mesh alternatives rather than create all 3D ground truth from scratch. Difficult edge cases go to skilled artists, and their corrections feed the next training round. This converts expensive geometry creation into a verification loop that can annotate far more real-world images while steadily improving the model that proposes each reconstruction.
The same general-purpose capability has a dual-use risk. Open concept detection, tracking, pose estimation, and 3D reconstruction can support accessibility, medicine, wildlife research, and robotics, but they also reduce the cost of comprehensive surveillance. Sidhu argues that builders need to evaluate deployment context and governance alongside technical performance.
Watch on YouTube


