SWE-2 (Fully Tested): WHAT? IT ACTUALLY BEATS ASTRA & FABLE!

AICodeKing9m 49s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    AICodeKing reports that SWE-2 scored 67 out of 80, or 83.75%, across eight KingBench 3 coding tasks. That narrowly beat DeepSeek V4.1 Flash in the same run and placed SWE-2 fifth on the channel's current leaderboard, above Fable 5 but behind GPT-6 Astra.

    The test spans interactive simulations, 3D objects, SVG generation, a game, a maths problem and a connected task that creates training data, fine-tunes a model and serves it through a local web interface. SWE-2 earned full marks on the maths and fine-tuning tasks, while DeepSeek V4.1 Flash performed better on the game and 3D wristwatch tasks.

    AICodeKing also found SWE-2 effective on longer agent work and a terminal-based movie tracker. Its main weakness was repeated clarification questions, so the suggested workaround is a system instruction that permits one essential question, encourages reasonable assumptions for routine choices and asks the model to continue toward a complete result.

    Original YouTube thumbnailWatch on YouTube