Noyan argues that sending every vision task to a large vision-language model is a poor fit for low-latency applications. Her proposed workflow uses a VLM to label images, smaller VLM judges to check the proposed annotations, and a task-specific detector such as RF-DETR for deployment. She treats model size, measured performance and licensing as separate selection questions.
The talk walks through road-sign detection and document-layout experiments. Noyan reports useful trained-model results but also finds strong disagreement between judges, requiring an explicit choice between stricter agreement and retaining enough training examples. Coding agents still need human-approved label descriptions and domain checks: flipping traffic signs or changing traffic-light colors can invalidate training data.
A second part packages specialist vision models as tools for coding agents, covering segmentation, detection, OCR and depth. Image-guided detection and stronger box-overlap checks are proposed future work, not demonstrated solutions to every industrial case.
Watch on YouTube




