Ari Morcos argues that better data can steepen a model's learning curve, giving more performance from a fixed compute budget or reaching a target with less computation. He frames quality as the information a model gains from each token and batch, emphasizing that the best dataset depends on the tasks the model needs to perform.
Ari Morcos describes four stages: clean, curate, create and compose. These include filtering unusable documents, removing benchmark contamination, reducing semantic redundancy, matching data to target tasks and balancing sources across training stages. He emphasizes relevance and diversity to make models less brittle when inputs change.
Ari Morcos reports results in vision-language and multilingual models where curated data improved measured capability and efficiency. He also describes shorter responses and cross-lingual transfer, including improvements in non-English performance after curating English data, while treating these as outcomes of particular experiments.
Ari Morcos explains synthetic rephrasing as converting information from selected source documents into multiple formats, such as true-or-false questions. He argues that choosing high-quality source documents matters more than rephrasing a random sample and that varied presentations help models learn more from useful material.
Ari Morcos discusses domain-specific mid-training that mixes specialized material with general training data to retain broader capabilities. In a legal-model example, he reports that this preparation also strengthened subsequent post-training gains, arguing that pre-training, mid-training and post-training should be designed together.
Watch on YouTube




