VentureBeat
Follow
85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one
Enterprises are increasingly removing humans from AI deployment decisions, even as trust in automated evaluation methods has risen. Despite this increased confidence, a significant portion of companies, around half, have experienced AI features that passed testing but then failed in production, a figure that has remained consistent. This indicates a growing gap between the perceived accuracy of evaluation layers and their actual effectiveness in preventing real-world failures. Companies that have experienced these production failures are actually accelerating their move towards automated deployments rather than slowing down. This counterintuitive trend suggests a focus on deployment maturity rather than a complete abandonment of automation after a test-passing failure. The data also reveals a lag in production quality monitoring, with many companies focusing on whether an agent functions rather than if its output is correct. The market for agent evaluation tools is evolving, with specialist platforms gaining traction and integration ease becoming a key purchasing factor. To manage the contradiction between increased autonomy and unreliable evaluations, enterprises are investing more in human-centered review workflows rather than solely relying on automated production observability. This creates a paradox where companies are removing human oversight for some deployment decisions while increasing investment in human review elsewhere in the process.