The AI agent evaluation gap: half of enterprises shipped agents that failed customers
VentureBeat Research survey of 157 enterprises reveals: 50% have shipped an AI agent that passed evals and failed in production. Only 5% fully trust automated evaluation, while 66% are moving toward zero-human-in-the-loop deployments.