All About AI proposes an ongoing prediction challenge between GPT-5.6 and Opus 5. The first round asks each model to estimate Seoul's lowest temperature and the number of dissenting votes at a Federal Reserve meeting, then places small positions on the selected outcomes so later videos can score the forecasts.
The benchmark tries to prevent the models from copying the market consensus. It rewrites each question to hide its Polymarket origin, removes Polymarket and Kalshi from search results, shows every possible outcome without prices and requires a probability distribution instead of a single yes-or-no answer.
Both models receive the same structured search interface and run one after the other in a cleaned repository. This keeps tool access comparable and reduces the chance that the second model will see the first model's output. One event per question also prevents several related market ranges from being counted as independent tests.
Opus 5 selected the market-favored 26-degree temperature range and two dissenting votes. GPT-5.6 chose 25 degrees and zero dissenters, producing more divergent positions for the first round. The video does not yet establish which model is more accurate because the underlying events had not resolved when it was recorded.
The creator treats this as an early version of the scaffolding and notes that Opus matching the leading indicators could mean the price-obfuscation method still leaks clues. Future rounds are intended to improve the controls, test more models and accumulate enough outcomes for meaningful comparison. The SerpApi sponsor segment is omitted from this summary.
Watch on YouTube



