Hiral Shah describes agreements as a large source of valuable but difficult-to-query business data. Individual contracts may contain pricing tiers, service levels, and other terms in tables with merged cells and nested layouts. Searching across many agreements also requires understanding which documents govern or modify others, so extracting plain text line by line can lose the structure needed for a reliable answer.
Sean Sodha separates the work into finding the right document and then finding the right information inside it. He describes a compact vision-language document parserA vision-language model jointly processes images and language so it can describe, answer questions about or act on visual information. that extracts text, reading order, layout, and table structure togetherIntelligent document processing uses AI to classify, extract and validate information from documents for downstream workflows.. The approach is designed for extraction rather than open-ended generation, and the speakers say a smaller purpose-built model helps keep processing cost and latency manageable at scale.
In a Docusign example, Hiral Shah walks through an uploaded order form whose terms and pricing table become structured fields that can be located in the agreement and exported for later analysis. The example shows the intended workflow, but the transcript alone does not verify every visual result or a general accuracy rate. The speakers later describe using a hybrid pipeline, including conventional OCR for some fields and layout-aware extraction for tables.
During audience questions, Hiral Shah and Sean Sodha distinguish bulk preprocessing from answering questions about one newly uploaded contract. A large corpus benefits from upfront extraction so later queries remain practical; a single-document question may justify a different latency and compute balance. They describe model compression and serving optimizations as future work rather than demonstrated improvements in this talk.
Watch on YouTube




