Scaling AI Legal Search Across Billions of Documents

AI Engineer20m 36s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Jacob Lauritzen distinguishes between searching documents inside a client's legal project and researching laws, cases, and regulations across jurisdictions. He explains how Legora first used shared Elasticsearch infrastructure, then regional clusters as data residency requirements grew. Enterprise customers also required stronger separation and control over encryption keys.

    Legora next put project search into PostgreSQL with pgvector, but Jacob Lauritzen says heavily partitioned indexes mixed frequently queried and idle projects, strained caches, and led to sharp latency spikes at scale. The team then moved to turbopuffer with one namespace per project. He reports better relevance through BM25 text search, lower operating complexity, and substantially improved latency. These results are vendor-reported within the talk, not independently benchmarked here.

    Simon Eskildsen explains turbopuffer's object-storage-first design: writes land in object storage, background work builds indexes, and queries prefer cached data before reaching storage. Namespace-level separation lets customers use distinct buckets and encryption keys. For some Legora workloads with stricter isolation requirements, the speakers say they disabled the disk cache and relied on memory plus object storage.

    For legal research, Jacob Lauritzen describes a corpus approaching ten billion vectors, queries that fan out across jurisdictions, and a mix of frequently accessed and rarely accessed law. Simon Eskildsen connects those access patterns to a memory hierarchy, clustered vector indexes, and compressed postings for full-text search. The technical trade-off is to keep hot data close to compute while leaving cold material in cheaper storage when a longer retrieval timeRetrieval is the process of selecting relevant stored information and returning it to an AI system for the current task. is acceptable.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Jacob Lauritzen in blue and Simon Eskildsen in dark clothing flank the white and blue headline LEGAL SEARCH AT SCALE on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 16 September 2026 and duration 20m 36s.

    Jacob Lauritzen and Simon Eskildsen trace Legora's move from shared search infrastructure to isolated, object-storage-backed retrieval for large legal AI workloads.