Data Quality Is the Compute Multiplier - Ari Morcos

AI Engineer19m 5s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Ari Morcos argues that better data can steepen a model's learning curve, giving more performance from a fixed compute budget or reaching a target with less computation. He frames quality as the information a model gains from each token and batch, emphasizing that the best dataset depends on the tasks the model needs to perform.

    Ari Morcos describes four stages: clean, curate, create and compose. These include filtering unusable documents, removing benchmark contamination, reducing semantic redundancy, matching data to target tasks and balancing sources across training stages. He emphasizes relevance and diversity to make models less brittle when inputs change.

    Ari Morcos reports results in vision-language and multilingual models where curated data improved measured capability and efficiency. He also describes shorter responses and cross-lingual transfer, including improvements in non-English performance after curating English data, while treating these as outcomes of particular experiments.

    Ari Morcos explains synthetic rephrasing as converting information from selected source documents into multiple formats, such as true-or-false questions. He argues that choosing high-quality source documents matters more than rephrasing a random sample and that varied presentations help models learn more from useful material.

    Ari Morcos discusses domain-specific mid-training that mixes specialized material with general training data to retain broader capabilities. In a legal-model example, he reports that this preparation also strengthened subsequent post-training gains, arguing that pre-training, mid-training and post-training should be designed together.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Ari Morcos smiling in a blue shirt beside “Data Quality Multiplies Compute” in blue and off-white on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 31 July 2026 and duration 19m 5s.

    Ari Morcos argues that task-relevant, diverse and carefully mixed training data can improve model performance per unit of compute, illustrating the approach with vision-language, multilingual, synthetic-data and domain-adaptation experiments.