Bijan Bowen tests Sakana AI's Fugu orchestration, which sends work across multiple models and combines their contributions. He compares Fugu and Fugu Ultra with individual models using the same practical coding and generation prompts.
The orchestrated systems complete several tasks successfully, but their extra calls and synthesis steps raise cost and latency. In multiple tests, the strongest standalone model matches or beats the combined result with a simpler execution path.
Bowen concludes that orchestration needs task-specific evidence rather than an assumption that more models automatically mean better answers. It may help on work that benefits from diverse attempts, but the overhead must be justified by measurable gains.
Watch on YouTube



