The Loophole framework turns a user's stated moral principles into a detailed rule system, then sends two adversarial agents looking for failures. One searches for immoral actions that remain legal under the generated rules, while the other finds moral actions that the rules prohibit. A judging agent patches straightforward translation errors and escalates genuine ambiguity or contradiction to the user.
The demonstration begins with DNA privacy and uses synthetic cases to expose missing boundaries in a moral code. The same pattern could test chatbot constitutions by searching for both forbidden answers that slip through and appropriate answers the system refuses, or compare personal data preferences with a company's terms and surface disagreements before a contract is accepted.
More speculative experiments model legislators and synthetic voter personas, then test or revise proposed bills against their inferred preferences. The presenter frames these government applications as aspirational and explicitly notes privacy and logistical problems, so the work is best understood as an exploratory stress-testing method rather than a validated representation of public values.
Watch on YouTube



