Arek Borucki on Scaling the Hugging Face Hub

AI Engineer21m 39s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Arek Borucki describes the Hub's metadata and application architecture, separating MongoDB records from model artifacts in object storage. This allows storage, metadata and compute to scale independently. The talk focuses on tail latency and keeping the user experience simple as the catalogue expands.

    Arek Borucki explains how model names are tokenized at insertion and stored in a denormalized read collection. The team replaced regular-expression search with Lucene-backed Atlas Search as the dataset grew, while maintaining ranking based on a changing popularity score.

    Arek Borucki describes distributing suitable reads and aggregations to replicas, reserving the primary for writes and consistency-sensitive operations, and isolating heavy reporting on a hidden node. Sharding is presented as a next step rather than an already completed migration.

    Arek Borucki distinguishes application-pod scaling from adding underlying cluster capacity. The talk also describes plans to move toward application-level demand signals rather than relying only on resource utilization. Reported scale figures belong to this presentation, not a continuously updated service benchmark.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Arek Borucki in a blue top beside the blue and white headline SCALING THE AI HUB on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 28 July 2026 and duration 21m 39s.

    Arek Borucki explains how Hugging Face separates model storage from metadata, optimizes search and distributes workloads to keep the Hub responsive as it grows.