Trail of Bits Blog
Follow
Auditing in the age of (good enough) AI
Security firms increasingly leverage AI beyond agentic code review to enhance security audits, building custom tools and formal models before actual code examination. For instance, in reviewing the Miden VM, which lacks developer tooling, agents spent six months creating an LSP server, a decompiler, a static analysis engine, and a Lean model of the VM executor. These tools identified significant security vulnerabilities, such as an unvalidated prover-supplied input that could enable forgery of Falcon signatures and fund theft. The Lean model also produced 95 machine-checked correctness proofs for a substantial portion of the Miden core library.The Miden VM, a new zero-knowledge VM with its own assembly language (MASM), presented unique review challenges due to its stack-machine architecture and absence of developer tools. To address this, the team prioritized building essential tools, starting with an LSP server and VS Code extension for syntax highlighting, code navigation, and instruction documentation. Next, a decompiler for MASM procedures was developed to provide high-level semantic information, despite complexities like implicit stack operations and un-declared signatures. This decompiler, a multi-month effort, yielded valuable internal analysis frameworks and an intermediate representation for static analysis.Utilizing this intermediate representation, the team built an abstract interpretation engine for static analysis, implementing passes to validate prover-supplied values, enforce type constraints, and ensure local variable initialization. This led to the discovery of over 400 potential type validation improvements and a high-severity bug: an underconstrained advice value in the mod_12289 procedure. This vulnerability allowed a malicious prover to manipulate remainders, potentially forging Falcon signatures and draining Miden accounts.Beyond bug finding, agents explored formal verification using Lean, modeling the Miden VM executor and automatically translating MASM procedures to Lean. This agent-driven formal modeling resulted in 95 correctness proofs for binary arithmetic components, uncovering two subtle bugs missed by existing unit tests: an edge case in 64-bit right-rotation (rotr) and an issue in 256-bit multiplication (wrapping_mul). These advanced tooling and formal modeling efforts, made feasible by recent advancements in AI agents, significantly improved the depth and quality of the security review, a level of preparatory work previously unattainable due to time and resource constraints.