All About AI reviews the first results from a comparison between GPT 5.6 and Claude Opus 5. Claude Opus 5 identifies the exact temperature outcome in one question, while neither model selects the resolved result in a second question. The small sample makes the exercise an early model-evaluation experiment rather than evidence of reliable forecasting performance.
The video then demonstrates an autonomous weather-analysis workflow that estimates expected value, opens and closes small prediction-market positions, and tracks performance across daily temperature questions. The presenter explains that early parameter choices produced losses and that the system is still being tested with conservative position sizes.
A separate research loop runs machine-learning experiments against a market baseline and records whether candidate strategies reduce loss on the observed data. The next stated step is evaluation on unseen data, which is necessary before the apparent improvement can be treated as a repeatable signal rather than overfitting or chance.
Watch on YouTube


