Brendan Rappazzo presents Morgan Stanley's AlphaLab, a provider-agnostic system for research tasks such as forecasting and model optimization. It combines contextual research, evaluation construction and repeated experiments rather than relying on a single agent to complete the whole process.
Research agents collect notes, evaluation builders work with conceptual and programmatic critics, and a strategist coordinates worker experiments through a Kanban-style queue and Slurm compute jobs. Human researchers can steer the queue and compare candidate models on held-out data. Reported benchmark and internal-model improvements are presented as the speaker's results, not independent validation or investment advice.
Rappazzo describes harness design and measurement errors as recurring problems. He argues that carefully constructed evaluation environments encode scarce enterprise expertise, support measurable iteration and can provide feedback for improving the harness itself. The presentation closes with plans for automated harness optimization and learning from successful research traces.
Watch on YouTube




