The Best Way To Test New AI Models

The AI Daily Brief46m 36s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Nathaniel Whittemore and Nufar Gaspar recommend selecting a small, varied set of recurring tasks and desired future workflows. Comparing a baseline with a few alternative models on identical inputs produces evidence relevant to one's actual work instead of relying on general benchmark rankings.

    Nathaniel Whittemore and Nufar Gaspar use fresh chats, concealed model labels and explicit scoring criteria to reduce familiarity bias. Scores should consider usefulness, quality, latency, cost and organizational constraints, with repeated trials where model variability matters.

    Nufar Gaspar demonstrates a personal evaluation spanning writing, negotiation, research, a web hub and planning tasks. Human judgments and automated judges can disagree, so an AI judge must be checked against the evaluator's own preferences rather than treated as an objective authority.

    Nathaniel Whittemore and Nufar Gaspar explain that the practical outcome may be staying with a familiar model, switching or allocating different tasks to different models. Sensitive work can use representative mock inputs, and subscription prices should not be confused with API costs.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Nathaniel Whittemore and Nufar Gaspar beside the blue and white headline YOUR OWN AI BENCHMARK on a black background. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 8 October 2026 and duration 46m 36s.

    Nathaniel Whittemore and Nufar Gaspar show how representative tasks, identical inputs and blind scoring can make model choices more useful than generic leaderboards.