VentureBeat
Follow
Amazon AGI director says AI agent reliability, not capability, is blocking enterprise deployment at VB Transform 2026
The enterprise AI industry faces a significant gap between piloting AI agents and deploying them in production. Bryan Silverthorn of Amazon attributes this to a flawed approach to evaluating AI agent reliability. He proposes breaking reliability into four dimensions: consistency, robustness, predictability, and safety. Current evaluations often fail to capture real-world failures, as demonstrated by an agent that intermittently read incorrect serial numbers due to subtle changes. Therefore, measurement rigor must match application stakes.Amazon's AGI lab manages AI agents like "interns," acknowledging their power and potential for error. This requires management skills, focusing on risk mitigation, backups, and undo capabilities. They accept occasional errors in exchange for faster research velocity. Silverthorn clarifies that fully autonomous self-improvement in AI is still a distant goal. AI agents will integrate with various tools for complex workflows. The key for enterprises to move beyond pilot phases is to prioritize consistent, correct performance over single impressive feats. Ultimately, successful AI agent deployment hinges on effective management rather than just sophisticated agents.