Outcome Intelligence: measuring AI by results, not activity
The first rung of Decision Assurance is to measure the real outcome instead of the agent's activity. It sounds obvious and it is rarely done, because activity is easy to measure and outcomes are not. This paper defines Outcome Intelligence: measuring each decision against the result it was meant to produce, connected back to it from the system of record or a person after the run. It gives a portable test for telling an activity metric from an outcome metric, explains why an outcome has to be defined as conditions before it can be measured, distinguishes an estimated outcome from a reconciled one, and argues for the honesty an outcome demands, including the discipline of saying "not yet reconciled" rather than inventing a success.
Outcome Intelligence measures each decision against the real result it was meant to produce, connected back from the system of record or a person, not read out of the agent's own account of its work. The portable test: if you can compute a metric from the agent's logs alone, it is activity, not outcome. Provy can estimate an outcome from those logs, but the reconciled verdict is not computable from them; it needs the settled result connected back and checked against the conditions set up front. Trust is built only from outcomes reconciled against reality, and where reality has not answered yet, the honest reading is "not yet reconciled," never a fabricated success.
- Outcome Intelligence measures each decision against the real result it was meant to produce, connected back from the system of record or a person after the run.
- Portable test: if you can compute a metric from the agent's own logs alone, it is an activity metric, not an outcome metric.
- An outcome has to be defined as conditions before it can be measured. That written definition is the Outcome Contract.
- An estimated outcome is the agent's own claim of success. A reconciled outcome is that claim checked against reality. Only the second is evidence.
- When the real outcome is not known yet, the honest answer is "not yet reconciled," never a fabricated success.
A health insurer ran an autonomous agent that adjudicated routine claims: it read the claim, checked the policy and the documentation, decided to approve, deny, or adjust, and closed the file. On its own record it was a strong performer. It cleared claims quickly, its decisions were well reasoned, and it closed nearly everything it touched. Then someone looked at the appeals queue. A meaningful share of the claims it had confidently closed were being reopened weeks later on appeal and overturned. The agent's dashboard still said those claims were adjudicated. The appeals system said they were adjudicated wrong. Two records of the same decisions, and only one of them was about the outcome.1
The agent was measuring itself the way almost every autonomous system does: by what it did. Claims read, decisions rendered, files closed. Those are real numbers and they were all good. None of them is the outcome. The outcome, whether the claim was adjudicated correctly and stayed adjudicated, showed up later, in a different system, and the agent never saw it. This paper is about closing that distance, which is the first and most fundamental commitment of Decision Assurance.
Activity is not achievement
The core confusion is between what an agent did and what it achieved. An agent doing things is not the same as an agent accomplishing the result those things were meant to produce, and the gap between the two is exactly where a confident-looking failure hides.
Outcome Intelligence is measuring each decision against the real result it was meant to produce, connected back from the system of record or a person after the run, rather than read out of the agent's own account of its work.
The reason activity gets measured instead of outcomes is not laziness. It is availability. The agent's activity is right there: it is emitted by the agent, in real time, in a form already built for counting. The outcome is somewhere else, arrives later, and often has to be understood before it can be scored. Faced with a metric that is free and a metric that is expensive, organizations reach for the free one and quietly let it stand in for the expensive one. That substitution is the whole problem. Claims closed is free. Claims that survived appeal is what mattered.
A test for telling them apart
There is a portable way to tell whether a metric is really about the outcome, and it is worth keeping because it survives being carried into any domain.
If you can compute a metric from the agent's own logs alone, it is an activity metric. A genuine outcome metric always requires a fact the agent's logs do not contain: something from the world, or the downstream system, or a later moment.
Run the claims agent through it. Claims processed, average handling time, decisions per hour, confidence scores: all computable from the agent's own record, all activity. Now try "share of adjudications overturned on appeal." You cannot compute it from the agent's logs at any level of detail, because the appeal happens in the appeals system, weeks later, and the agent's record closed long before. That single metric requires a fact from outside the agent, which is precisely what makes it an outcome metric and precisely why the agent could look excellent while being wrong. The test is blunt and it is reliable: the moment a number needs something the agent could not have known, you have left activity and reached an outcome.
This also explains why you cannot get outcome intelligence by adding more instrumentation to the agent. More logs, more spans, more self-reported scores are all more activity. The outcome is not under-instrumented; it is elsewhere. Reaching it means having the settled result connected back to you, not watching the agent more closely.
Defining the outcome as conditions
Before an outcome can be measured, it has to be defined, and vaguely is not good enough. "The claim was handled well" cannot be checked. What can be checked is a set of conditions: the decision matched the policy, the documentation supported it, the payout was correct, and it was not overturned on appeal within the appeal window. A decision that meets all of them achieved the outcome. A decision that misses one did not, and you know exactly which one.
Writing those conditions down, before the agent runs, is the discipline captured in the Outcome Contract: the agreed statement of what a good result looks like, expressed as conditions a real outcome can be checked against. It matters that this is a list rather than a single verdict. Grading a decision as one pass-or-fail throws away the information that makes a failure fixable. "Three of four conditions met, the appeal condition missed" tells you where to look. "Failed" tells you nothing. Defining the outcome as conditions is what turns a later measurement from an opinion into a check.
This is also the point where the two owners of Decision Assurance meet. The business owns the conditions, because only the people who run claims can say that surviving appeal is part of what "adjudicated correctly" means. The platform owns making those conditions measurable against the real result. Outcome Intelligence is where a written condition becomes a measured fact.
Estimated versus reconciled
Not every claim of success is the same kind of claim, and conflating two of them is one of the most common mistakes in operating autonomous AI. There is the outcome the agent believes it achieved, and there is the outcome reality confirms. They are different objects and only one of them is evidence.
When the claims agent closes a file marked "approved, correct," that is an estimated outcome. It is the agent's best judgment at the moment of decision, and it is useful: it is how the agent knows what it was trying to do. But it is a prediction, not a result, and it carries exactly the confidence problem this whole library is about, because a confident agent estimates success even when it is wrong. A reconciled outcome is that estimate checked against the settled result connected back to Provy a month later: did the claim actually stay approved, or was it overturned? The two agree most of the time. The entire value of the discipline is in the times they do not, because that is where a silent failure is caught. This distinction has its own concept, Estimated vs Reconciled, because so much depends on never mistaking the first for the second. A dashboard built on estimated outcomes is just activity wearing the word "outcome." Only reconciliation makes it real.
Provy computes the estimated outcome from the run's own logs, the system's own read at the moment it acted. The reconciled verdict is not computable from logs at any depth. It waits for the settled result to be connected back to Provy, tagged to the same work item, and checked against the conditions written before the run. Estimated is derived. Reconciled is supplied.
The honesty of outcomes
Measuring real outcomes forces an honesty that measuring activity lets you avoid, and it is worth being direct about what that honesty costs and requires.
Outcomes are delayed. The appeal arrives weeks after the decision. For that window the outcome is genuinely unknown, and no amount of looking at the agent will resolve it early. A system that measures outcomes has to be able to hold a decision open, unresolved, until reality answers, rather than closing it as a success the moment the agent finishes.
Outcomes are partial. A claim can be right on the payout and wrong on the coding, approved correctly but adjudicated under the wrong provision. Defining the outcome as conditions is what lets a partial result be recorded as what it is, some conditions met and others not, instead of being flattened into a success or a failure that hides the truth.
Outcomes have coverage limits. Some decisions can be reconciled cleanly, others cannot yet, because the downstream signal is not connected or has not arrived. An honest outcome measurement reports its own coverage: how many of the decisions have actually been checked against reality, and how many are still riding on the agent's estimate. A high success rate over a thin slice of reconciled outcomes is not the same as a high success rate, and pretending otherwise is how a number lies.
When the real outcome is not yet known, the honest answer is "not yet reconciled." It is never a fabricated success. A system that fills an unknown outcome with an optimistic guess has quietly gone back to measuring activity, and has done it while claiming to measure results.
All of this raises a genuinely hard question, and it is worth naming it precisely, because the discipline is defined in part by taking it seriously. When a reconciled outcome comes back wrong, one of two very different things happened. Either the agent made a bad decision, or the agent made a reasonable decision and the world moved underneath it: the policy changed, a fact that was true at decision time stopped being true, the appeal turned on evidence that did not exist when the claim was closed. Was the agent wrong, or did the world move? Answering it fairly is what separates a real accounting of an autonomous system from a blame machine that punishes agents for being caught in a changing world. This library names the question as one the discipline must answer; how it is answered is beyond its scope.2
In practice
One idea, seen across different kinds of work. These are illustrations of the concept, not product walkthroughs.
Resolved, or just closed? A support agent marks a case resolved and its own confidence is high. That is an estimated outcome. The reconciled outcome arrives when the customer either stays quiet or reopens the case a week later, and only the second belongs on a trust dashboard.
Matched, then signed off. A reconciliation agent reports a batch fully matched. Estimated. Whether those matches survive the period-close review, when a controller signs off, is the reconciled outcome, and the distance between the two is the overconfidence to watch.
High intent is not a closed deal. A lead-scoring agent marks an opportunity high-intent and it reads like a win. Estimated. Whether the deal actually closed, weeks later in the CRM, is the reconciled outcome, and a pipeline built on the first number flatters itself.
Outcome Intelligence is the first rung of the ladder for a reason. Everything above it depends on measuring the real result: there is no evidence worth holding and nothing worth verifying if the thing you are proving and re-checking is the agent's own account of its work. Measure the outcome first. The next paper takes up what you have to keep so you can stand behind that measurement later, which is Runtime Evidence.
Concepts in this paper
Notes
- The example is illustrative. The pattern, a decision that looks correct at the moment it is made and is revealed as wrong only by a later, downstream signal, generalizes across claims, finance, support, and sales.
- The question "was the agent wrong, or did the world move?" is named here as a problem the discipline must take seriously. This library does not describe how it is answered.
- "Coverage" is used to mean the share of decisions that have actually been checked against a real outcome, as distinct from the share the agent estimated as successful.
References
- Provy Research. Decision Assurance: operating autonomous AI with confidence. Provy Knowledge Center, 2026.
- Provy Research. Runtime Evidence. Provy Knowledge Center, 2026.
- Provy Research. The Enterprise Confidence Gap. Provy Knowledge Center, 2026.
Change history
- v2.2 · July 2026 · Made the mechanism honest: Provy computes the estimated outcome from the run's logs, but the reconciled verdict is not computable from logs. It needs the settled result connected back and checked against conditions written up front. Corrected phrasing that implied Provy reads the downstream system.
- v2.1 · July 2026 · Added the publication number, an executive summary, enterprise scenarios, and the signature diagram.
- v2.0 · July 2026 · Rebuilt as part of the numbered research series.
How to cite
Measure your agents by the outcome, not the activity
Provy takes the settled result you connect back to it and checks every decision against the conditions you set up front, so a confident-looking failure has nowhere to hide.
Get a demo