Enterprises are running dozens of agent pilots, and most of them will die this year. Not because the model hallucinated, not because latency was bad — because when the budget review came, nobody could prove the pilot was used.
Walk through any large enterprise and count the agent pilots: a support copilot, an HR bot, a SQL assistant, two coding agents, something in procurement nobody fully remembers approving. Each one launched with a demo that impressed a steering committee. Twelve months later, most are gone.
The post-mortems rarely say "the model wasn't good enough." They say things like "unclear value," "couldn't quantify impact," "sponsor moved on." Translate those, and they all mean the same thing: when someone finally asked "is anyone actually using this, and what does a successful outcome cost?", the room went quiet.
A demo answers "can this work?" — and modern models answer that question so convincingly that pilots get approved in a single meeting. But the question that decides renewal is different: "did this work, for whom, how often, and at what cost per outcome?" That question can't be answered with a demo, an anecdote, or a screenshot of a great conversation. It needs boring, continuous, comparable numbers:
Notice what's not on the list: eval scores, token counts, latency percentiles. Those live in your tracing tool and matter enormously — to engineers. A steering committee cannot read a trace, and it shouldn't have to. (We wrote a full guide on this split: how to measure AI agent pilots.)
Three reasons, all fixable:
You don't need an analytics team per pilot. You need four events from the agent's backend — conversation started, capability used (with its cost in USD), an outcome (task-completed, task-failed, or escalated-to-human), and a user verdict. Wire those into any measurement layer and every pilot in the company becomes comparable on one table, sorted by cost per completed task. The fund/fix/retire conversation mostly runs itself from there.
If a pilot can't survive honest adoption numbers, it shouldn't survive the budget review — and that's the system working. The point of evidence isn't to save every pilot. It's to make sure the ones that die deserve it, and the ones that scale can prove they do. In 2026, "we believe it's helping" is not a line item. Bring a number or bring your farewell slide.