Knowledge Center/ Insights/ Did versus achieved
Insight

You are measuring what your agents did, not what they achieved

Two habits keep teams grading the wrong thing: counting activity because it is easy, and treating a compound success as a single check. Both have a simple fix.

Provy · Insight 6 min read July 2026

Most teams running autonomous agents can tell you, to the decimal, how much their agents did. Tasks handled. Tickets closed. Approvals issued. Follow-ups sent. What far fewer can tell you is whether any of it worked, and the gap between those two questions is where confidence quietly leaks. The trouble is that the first question is easy and the second is hard, so teams answer the easy one and quietly hope it stands in for the hard one. It does not, and there is a clean test for spotting when you have made the substitution.

The razor: who could compute this number

Here is the test. Take any metric you report about your agents and ask a single question: could I compute this from the agent's own logs alone? If the answer is yes, it is an activity metric, no matter how much its name sounds like an outcome.

Run it on the usual suspects. "Resolution rate" is computed from the agent marking tickets resolved. "Approval rate" is computed from the agent recording approvals. "Follow-ups completed" is computed from the agent logging that it sent them. Every one of these is calculated entirely from the agent's record of its own behavior. The agent is both the actor and the scorekeeper, and no scorekeeper who is also the player is measuring the result. It is measuring the effort.

If you can compute the number from the agent's own logs, you are measuring what it did. An outcome needs a second witness the agent cannot write.

A real business outcome fails the test, and that is the point. To know whether a sales follow-up worked, you need what happened in the deal: did it advance, or did it die somewhere the agent never sees. To know whether a financial reconciliation worked, you need whether the ledger tied out at close, not whether the agent reported that it did. In both cases the verdict comes from a source outside the agent, arriving later, that the agent has no way to fake because it does not control it. That second witness is the whole difference between activity and achievement. Outcome Intelligence is the discipline of getting that verdict connected back into the loop.

How "achieved" becomes known

Provy does not reach into your ledger, your ticket queue, or your CRM and read them. Nothing divines the result, Provy included. The settled outcome is connected back to Provy after the run, keyed to the same work item, by the system of record or by a person. Provy then checks it against the conditions you wrote in advance. Until a result is connected, the honest verdict is "not measured," never a quiet pass.

The reason this matters is not tidiness. An agent optimized against an activity metric will get better at the activity, which is not the same as getting better at the job. Reward "tickets closed" and you will get more closed tickets, some of which are closed on customers whose problem is still live. The metric goes up. The outcome goes down. You cannot see it, because the only number you are watching is one the agent gets to write.

Success is several conditions wearing one coat

Suppose you fix the first problem and start measuring the real outcome. There is a second trap waiting, subtler than the first. Even a genuine outcome usually hides a conjunction inside a single word.

Take that reconciliation again. What does "done right" mean? The totals tie out. The entries land in the correct period. No transaction is counted twice. Anything that could not be matched is flagged rather than silently dropped. And it finished before the close. That is five conditions, and the run is a success only if all five hold. But the natural way to track it is one green check labeled reconciled, which collapses the five into one and hides four of them. A run that ties out but posts to the wrong period satisfies the headline and fails the job, and the single check has no way to show it.

This is why "success rate" as a lone percentage is so often a comforting fiction. It is the average of a compound nobody decomposed. The fix is not a better number. It is to stop hiding the conjunction: write down what a good run requires, as a list of separate conditions, before the run, and grade each one. Not "did it succeed" but "how many of the five conditions held, and which one broke." A run that meets four of five is a specific, actionable thing. A run that is simply "80%" is a shrug.

Naming the conditions in advance is what turns a vague sense of success into something you can grade and defend. That commitment has a name, the outcome contract, and its only real requirement is that you decide what done right means before you can be tempted to grade against whatever the agent happened to do.

Two habits, one correction

Both traps share a root. In each, the agent's own account of its work has been allowed to stand in for the result. The activity metric trusts the agent to score itself. The single green check trusts a label to summarize a conjunction it never checked. Correct both the same way. Measure against a witness the agent does not control, and decide in advance what the witness has to show. Do that, and you are finally measuring what your agents achieved instead of what they did.