Gemini Wins the AI Chess Tournament

No Hype AI18:32
0 comments · 0 votesOpen discussion
Sign in to join the discussion

    Video summary

    The experiment treats chess as a demanding reasoning benchmark that avoids many familiar evaluation problems. Head-to-head games cannot be memorized because each position is novel, and the enormous search space leaves ample room above today's language-model performance, unlike benchmarks that quickly saturate or leak into training data.

    Eight proprietary and open models first solve Lichess puzzles using FEN positions. GPT and Gemini score 19 out of 20 on the easy set, while Claude, Grok and the open models trail. Adding an ASCII board changes little, and supplying every legal move makes results worse because some models spend their reasoning budget exploring the longer option list instead of solving the position.

    The top four then play a tournament. GPT advances past Grok despite repeated tactical mistakes, while Gemini beats Claude twice by exploiting hanging pieces and converting decisive advantages. Gemini then defeats GPT 2-0 in the final, showing more consistent board awareness, tactical conversion and checkmating ability than its opponents.

    On harder puzzles, Gemini scores 44 out of 50 in the medium band and 16 out of 50 in the hard band, producing an estimated Glicko puzzle rating of 2,087. The presenter cautions that puzzle ratings exceed practical playing strength and that the public Lichess database may introduce some contamination, but argues that the same test can still track model improvement over time.

    Original YouTube thumbnailWatch on YouTube