Claude Sonnet 5 Put to the Test

Pat Simmons22:53
1 VIEW
0 comments · 0 votesOpen discussion
Sign in to join the discussion

    Video summary

    Pat Simmons evaluates Claude Sonnet 5 against Opus 4.8 and Sonnet 4.6 with blind side-by-side tests rather than relying only on published benchmark scores. The coding set covers a designer portfolio, an Excalidraw-style drawing tool, a particle galaxy, a racing game and an endless terrain flyover.

    The coding results are mixed. Sonnet 5 closely matches Opus on the simpler portfolio task at a lower cost, but several more demanding interactive builds fail or underperform. Opus is the most consistent coding model in this small test set, while the older Sonnet 4.6 is often weaker and can cost more because it uses additional output tokens.

    The knowledge-work section compares a sales presentation, essay completion, a LinkedIn post and a YouTube hook. Sonnet 5 produces the preferred presentation design and the preferred concise video hook, while Opus is preferred for the longer essay. Simmons stresses that writing judgments are subjective and heavily affected by prompts, examples and feedback.

    His conclusion is that Sonnet 5 can be a cost-effective choice for landing pages and common knowledge work, but it does not replace Opus for every task. The tests are informal demonstrations rather than a controlled evaluation, and model pricing and promotional discounts are time-sensitive. The creator's bootcamp and newsletter promotions are omitted.

    Original YouTube thumbnailWatch on YouTube