Nathaniel Whittemore and Nufar Gaspar recommend selecting a small, varied set of recurring tasks and desired future workflows. Comparing a baseline with a few alternative models on identical inputs produces evidence relevant to one's actual work instead of relying on general benchmark rankings.
Nathaniel Whittemore and Nufar Gaspar use fresh chats, concealed model labels and explicit scoring criteria to reduce familiarity bias. Scores should consider usefulness, quality, latency, cost and organizational constraints, with repeated trials where model variability matters.
Nufar Gaspar demonstrates a personal evaluation spanning writing, negotiation, research, a web hub and planning tasks. Human judgments and automated judges can disagree, so an AI judge must be checked against the evaluator's own preferences rather than treated as an objective authority.
Nathaniel Whittemore and Nufar Gaspar explain that the practical outcome may be staying with a familiar model, switching or allocating different tasks to different models. Sensitive work can use representative mock inputs, and subscription prices should not be confused with API costs.
Watch on YouTube




