Maximilian-David Rumpf argues that current agent-driven search can improve document finding but may be slow and costly. Conventional pipelines chain query rewriting, retrievalRetrieval is the process of selecting relevant stored information and returning it to an AI system for the current task. and reranking with fixed decisions; a reranker that sees weak results cannot independently reformulate the search.
The proposed reinforcement learningReinforcement learning trains an AI system to choose actions using feedback about the results of its earlier choices. approach gives a specialist model repeated access to a search database so it can inspect results, revise queries and filters, and return a ranked listRanking orders candidate items by a defined score or criterion so an AI application can prioritize the most relevant or useful results.. The talk frames correct-document retrieval as a measurable training reward and argues that specialization can use compute more efficiently than a general model.
Maximilian-David Rumpf presents a mixed academic and internal benchmark in which his team's search model is reported to run about 20 times faster and cost about 100 times less than a frontier-model search approach. These are speaker-reported comparisons, not independently verified guarantees. He also proposes a search subagent that filters poor results before passing evidence to a main agent, reducing irrelevant material in its context.
Watch on YouTube




