From Raw Documents to AI-Ready Data - Leo Platzer & Jeff Koss

AI Engineer21m 37s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Jeff Koss and Leo Platzer use a manufacturing-company scenario to contrast a successful chatbot pilot over selected files with failures after expanding to a much larger document collection. Stale contracts, duplicates, sensitive material and irrelevant documents can undermine answers even when the model and retrieval pipeline appear to work.

    Jeff Koss demonstrates document classification, evidence-backed metadata and filters for relevance, sensitivity and freshness. Leo Platzer explains why selecting the right data slice matters: repeated or outdated documents can crowd useful evidence out of retrieval results. Classification confidence still needs review and is not a guarantee of correctness.

    Leo Platzer and Jeff Koss describe maintaining curated collections through an SDK and making context available to agents in repository-friendly files. They present retrieval improvements from their own experiments, while the broader lesson is to evaluate data preparation against the actual task rather than treating more documents as automatically better.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Leo Platzer and Jeff Koss against a black background beside the blue-and-white headline “MAKE DOCUMENTS AI READY”. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 5 October 2026 and duration 21m 37s.

    Leo Platzer and Jeff Koss demonstrate how governed metadata and quality filters can turn sprawling document collections into more useful AI retrieval inputs.