VentureBeat
Follow
The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway
Organizations are increasingly granting AI agents more autonomy, yet they are losing trust in the evaluations designed to control that autonomy. A significant fifty percent of companies have deployed an AI agent that successfully passed internal evaluations but subsequently failed with customers in production. Currently, only a meager five percent of organizations fully trust their automated evaluation processes. The primary identified weakness is that these evaluations do not accurately reflect real-world outcomes. Despite this, a substantial two-thirds of companies already permit, or are developing systems to allow, the deployment of agent changes directly to production based solely on automated evaluations, without human oversight. This disparity creates an "evaluation gap," signifying the difference between the autonomy granted to agents and the insufficient trust in the tests meant to monitor them. The research examines how leaders measure agent performance, the platforms they employ, and their willingness to allow unsupervised agent operation. Half of organizations have experienced customer-facing failures from agents that passed internal checks, and a quarter have seen this happen multiple times. Only five percent fully trust automated evaluations, primarily due to poor alignment with real-world results. Nevertheless, sixty-six percent of organizations are moving towards or already permit zero-human-in-the-loop deployments for agents. The evaluation and reliability tooling landscape is fragmented, with provider-native tools and "no dedicated tooling" being the most common. Furthermore, only about a quarter of companies conduct real-time quality checks on live production traffic, leaving a significant blind spot in monitoring agent output correctness. Enterprises select evaluation tooling based on cost and integration, with consistency being the key measure of success. Future investment is anticipated to increase for both human oversight and observability of AI agents.