Skill issue: stop deploying vision language models, use them with Skills - Merve Noyan, Hugging Face

AI Engineer19m 13s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Noyan argues that sending every vision task to a large vision-language model is a poor fit for low-latency applications. Her proposed workflow uses a VLM to label images, smaller VLM judges to check the proposed annotations, and a task-specific detector such as RF-DETR for deployment. She treats model size, measured performance and licensing as separate selection questions.

    The talk walks through road-sign detection and document-layout experiments. Noyan reports useful trained-model results but also finds strong disagreement between judges, requiring an explicit choice between stricter agreement and retaining enough training examples. Coding agents still need human-approved label descriptions and domain checks: flipping traffic signs or changing traffic-light colors can invalidate training data.

    A second part packages specialist vision models as tools for coding agents, covering segmentation, detection, OCR and depth. Image-guided detection and stronger box-overlap checks are proposed future work, not demonstrated solutions to every industrial case.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Merve Noyan against a black background beside the blue and white headline “TRAIN SMALL / SEE FAST”. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 23 September 2026 and duration 19m 13s.

    Use vision-language models to label and judge data, then deploy a small specialist model. Merve Noyan explains the pipeline, its promising tests and the human checks coding agents still need.