Sean Cai distinguishes authentic professional trajectories from manufactured training tasks. Finished spreadsheets and records omit the decisions, tool use, recovery and changing constraints that produced them, so collecting more examples is not automatically the same as improving training signal.
He evaluates domains through decomposition, expert agreement and access to verified real examples. Finance-task illustrations show how similar aggregate scores can conceal opposite arithmetic and methodological weaknesses. A benchmark also measures its scaffold and grading conditions, and vendors selling both tests and training material can create misleading incentives.
Cai describes opportunities in enterprise model routing, repeated post-training and maintained reinforcement-learning environments. Robotics provides a counterexample because the right data modality remains unsettled. Market estimates, vendor examples and historical comparisons are the speaker's analysis rather than independently verified forecasts.
Watch on YouTube




