Theo Browne argues that a single AI model tier listA model leaderboard ranks AI systems using scores from a defined benchmark or collection of evaluations. is inherently reductive because model quality depends on the task. He compares models using code quality, output-token efficiencyAI token efficiency measures how effectively a model or workflow turns consumed input and output tokens into useful results., speed, cost, vision support and how reliably they follow instructionsAI instruction following is a model's ability to understand and reliably comply with a user's valid requirements and constraints. in real development workflowsWorkflow-specific AI evaluation tests systems using the tasks, tools, constraints, and outcomes of a real working process..
Theo Browne places Fable 5 alone at the top for code he is willing to merge, with OpenAI 5.6 Soul close behind as a more controllable and token-efficient general workhorse. He rates OpenAI 5.6 Luna highly for inexpensive structured tasks, while Kimi K3 and DeepSeek V4 Flash remain strong open-weight options with different tradeoffs.
Theo Browne is more critical of expensive or inefficient models, particularly those without vision or with unpredictable reasoning costs. His final ranking favors models that produce useful work with fewer tokens and less supervisionAI supervision burden is the human time and effort required to direct, inspect, correct and approve AI-generated work., rather than those that lead on isolated benchmarks or raw generation speed.
Watch on YouTube



