Model Evaluation Videos

Videos that compare AI models through benchmarks, hands-on tests, cost analysis, and practical task performance.

Search the index

Find videos

Showing 41–60 of 157 videos

Clear filters
  1. Nathaniel Whittemore beside the headline Fable 5.1 Earns Its Place
  2. A flat illustrated game world beside the words Fable 5.1 Builds Whole Worlds
  3. Flat benchmark cards beside the headline AI Models Are Hard to Test
  4. Nick Saraev and Jack Roberts discuss frontier AI costs and safety
  5. A flat benchmark chart beside the words Fable 5.1 Wins the Bench
  6. Pat Simmons beside the headline Fable 5.1 Cheaper Not Better?
    How Fable 5.1 Compares in Hands-On Tests
    Pat Simmons29m 59s2 VIEWS
  7. Matthew Berman beside the headline Fable 5.1 Cuts Agent Costs
  8. Bijan Bowen beside the headline Fable 5.1 Builds Big Spends Big
  9. Wes Roth beside the words Fable 5.1 Beats Astra
  10. Alex Finn beside the words Fable 5.1 Gets Leaner
    Claude Fable 5.1 Gets Faster and Cheaper
    Alex Finn12m 27s1 VIEW
  11. Theo Browne beside the words Reward Hacking Spreads
  12. Bijan Bowen beside the words HY4 Preview Tested
  13. Kanish Manuja beside the words LLM Gateways That Survive Production
  14. Nachiket Paranjape and Swaroop Chitlur Haridas beside the words AI Quality Is a Team Sport
  15. Ryan Greenblatt beside the words AI Agents Learned to Cheat
  16. Bijan Bowen in a pink T-shirt beside FLASH in white and PUNCHES UP in pink on a black background.
  17. The words Agents Broke the Boundary beside a flat security shield illustration
  18. Simran Arora beside the words LLMs Struggle With Multi-GPU Kernels
  19. Theo Browne examining coding-agent behavior and pull-request reviews with GLM 5.3 Flash
    Ox Alpha is INSANE
    Theo43m 15s
  20. Bijan Bowen beside the words Qwen 4 Preview Impresses