Filip Makraduli starts with an Obsidian vault containing sensitive documents and a desire to reduce context costs. His workflow sends document-processing jobs through an MCP serverModel Context Protocol is a standard way for AI applications to connect with external tools and data sources through consistent interfaces. to an inference clusterAI inference is the process of running a trained model on new input to produce a prediction, classification, generated response or action. under the team's control, then returns a smaller Markdown artifact to the main coding agentAn AI coding agent is a tool-using AI system that can inspect, modify, and validate software within a repository.. The demonstration uses synthetic records rather than real confidential files.
Filip Makraduli shows OCR and entity-redaction tools processing an NDA and a scanned contract. The intended boundary is that the downstream agent receives the redacted output, not the original document. He distinguishes the managed cluster used for the demonstration from infrastructure a team could self-host, and acknowledges that privacy guaranteesData privacy governs how personal, confidential, or sensitive information is collected, used, shared, retained, and protected in AI systems. need further testing and guardrails.
Filip Makraduli connects this workflow to inference infrastructureAI compute infrastructure is the hardware, facilities, networks, power, cooling, storage, and software used to train and run AI models.: per-worker batching, warm models, LoRA swapping, hardware choices and scaling. He reports internal embedding comparisons but does not establish a universal cost or accuracy advantage. Task-aware model routing and councils of models are described as future directions rather than completed features, with reasoning budgets and biased evaluators among the risks.
Watch on YouTube




