Model Evaluation Videos

Videos that compare AI models through benchmarks, hands-on tests, cost analysis, and practical task performance.

Search the index

Find videos

Showing 121–140 of 157 videos

Clear filters
  1. Portrait of Bijan Bowen beside the words Two Agents Build the Game
  2. Portrait of Bijan Bowen beside the words Small Model Big Coding
  3. Alejandro Vidal beside the words Benchmarks Miss Difficulty
  4. Susheem Koul and Tisha Chawla beside the words Replay Agent Failures
  5. Portrait of Nate B. Jones beside the words Test the Task Not the Flag
  6. Theo Browne beside the words Open Model Closes the Gap
  7. The words Opus 5 Pushes Harder beside a single dark monolith
    OPUS 5 is Actually INSANE
    AI Copium11m 35s
  8. Portrait of Pat Simmons beside the words Opus 5 Wins the Tests
  9. Portrait of Bijan Bowen beside the words Opus 5 Builds with Detail
  10. Portrait of Alex Finn beside the words Opus 5 Needs Focus
  11. Portrait of Bijan Bowen beside the words Laguna Writes Better Than It Codes
  12. Portrait of Pat Simmons beside the words Open Model Closes the Gap
  13. Bijan Bowen weighing two options beside the words Real Tests Pick a Winner
  14. Bijan Bowen beside the words GLM 5.2 vs Claude
  15. Bijan Bowen beside the words Ornith 1.0 Goes Local
  16. Bijan Bowen beside the words Fugu Team Costs More
  17. Nate B. Jones beside the headline Pick Models by the Work
  18. Portrait of Bijan Bowen beside the words Gemini Flash Is Fast but Uneven
  19. Portrait of Pat Simmons beside the words Kimi Wins the Hard Builds
  20. Portrait of Bijan Bowen beside the words Qwen Max Shows Promise