What Ox Alpha's Benchmarks Reveal

AICodeKing5:36
0 comments · 0 votesOpen discussion
Sign in to join the discussion

    Video summary

    Ox Alpha appeared anonymously through OpenCode with temporary free access, a claimed one-million-token context window, multimodal input and zero data retention. AICodeKing tested the model on King Bench and says it scored 70 out of 80, placing second on that leaderboard behind GLM 5.3.

    The model earned full marks on several web, SVG, mathematics and fine-tuning tasks, while losing points on a folding-table task and a difficult 3D clock test. A separate ten-task subset reported by Ben Davis put Ox Alpha ahead of the comparison models, although the small sample means the result could have high variance.

    The video's provider attribution remains an informed hypothesis rather than a confirmed fact. Davis's investigation found matching video-token behavior, tokenizer counts and response patterns between Ox Alpha and models in the GLM family, while other candidate providers showed different fingerprints.

    AICodeKing concludes that Ox Alpha may be tuned more strongly for agentic coding than one-shot generation and that its early results deserve further testing. The summary omits the channel's closing requests for memberships, donations, likes and subscriptions.

    Original YouTube thumbnailWatch on YouTube