David Ondrej distinguishes mixture of agents from mixture of experts. In this workflow, several complete reference models answer the same prompt independently, then a separate aggregator combines their proposals and decides the final action. Hermes exposes each configuration as a selectable model while retaining tools, memory and session context.
The walkthrough creates a preset with GLM 5.2, GPT 5.5, Kimi K2.7 Code and Opus 4.8 as references, with Opus 4.8 as the aggregator. Ondrej recommends higher reference temperature for diverse proposals and lower aggregator temperature for a more consistent decision, while noting that different presets can target coding, research or review.
A second agent monitors Hermes in another terminal pane, sleeps between checks and steers the run when the aggregator appears stuck. The mixture builds and deploys a working 3D game, but the process takes about twenty minutes and roughly twenty dollars, demonstrating that it is unsuitable for quick or routine questions.
The practical takeaway is to reserve mixture of agents for expensive problems where independent viewpoints can improve planning, debugging or review. More models mean more tokens, longer waits and more failure points, so observability, spending limits and active supervision remain essential. A lengthy Hostinger sponsorship and repeated promotions for downloadable resources, skills and social accounts are omitted from the catalogue record.
Watch on YouTube


