Rashi Agrawal presents three layers for member-facing health AI: privacy constraints built into the architecture, deterministic controls above the language model and continuous safety evaluation. Drawing on work at Hinge Health, Rashi Agrawal describes removing protected health information at ingestion, separating production from nonproduction environments and restricting access by role and geography. These are proposed engineering safeguards, with no independent assessment of their effectiveness presented in the talk.
Rashi Agrawal argues that prompts should not carry the burden of enforcing security boundaries. In the proposed stack, code runs before the model on every conversation turn and controls emergency escalation, routing into clinical or support capabilities, and identity checks before access to member data. A model can assist with classification and conversation, while the surrounding software owns consequential routing and authorization decisions.
Rashi Agrawal describes ongoing evaluation through automated judges, member feedback and human review of sampled conversation traces, with complete review proposed for high-stakes cases. New failure patterns should lead to additional monitoring, because prompt fixes can stop working as prompts, tools and models change. Rashi Agrawal identifies the availability of people who can interpret and act on those signals as a practical constraint.
Rashi Agrawal offers a launch decision framework that assigns severity according to the worst plausible harm, keeps severity separate from engineering capacity and favors delaying when an unresolved issue concerns safety. The framework also compares proposed launch standards with risks already accepted in production, requires explicit decisions to fix, delay or accept a risk, and treats promised follow-up fixes as commitments. These are Rashi Agrawal's operating principles rather than evidence that existing production behavior establishes clinical safety.
Rashi Agrawal closes by examining errors in the evaluators themselves. A declining automated score can reflect a faulty judge or a faulty agent, so reviewers should inspect the underlying case before changing the system. Rashi Agrawal uses contrasting examples to show why evaluator prompts need maintenance alongside the AI application; the examples illustrate an evaluation method rather than provide medical advice.
Watch on YouTube



