Stop Chunking Like It's 2022 - Yuval Belfer, AI21 Labs

AI Engineer18m 1s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Yuval Belfer argues that agentic search has not eliminated the need to prepare data for retrievalRetrieval is the process of selecting relevant stored information and returning it to an AI system for the current task.. Conventional RAG systems choose chunk sizes during indexing, before they know the future queries, so very large chunks can blur useful details while very small chunks can lose wider context.

    Yuval Belfer describes tests that index the same material at multiple chunk sizes across meeting-transcript, narrative-question-answering and other datasets. Focused and context-heavy questions favor different chunk sizes, and an oracle comparison illustrates the recall left on the table by committing to one fixed size.

    Yuval Belfer proposes multi-scale indexing: make several differently chunked indexes, retrieve from each, promote matches to document-level rankingsRanking orders candidate items by a defined score or criterion so an AI application can prioritize the most relevant or useful results. and merge those rankings with reciprocal rank fusion. He reports stronger recall across the tested datasets than individual fixed-size baselines, while acknowledging extra memory use and open questions about how many sizes to retain and whether a better fusion method exists.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Yuval Belfer in a black jacket on black beside the blue-and-white headline STOP FIXED CHUNKING. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 16 September 2026 and duration 18m 1s.

    Yuval Belfer argues fixed-size chunks miss query-specific evidence and reports that multi-scale indexing with reciprocal rank fusion improves RAG recall at a memory cost.