Knowledge Center/ Research/ The Honest Limits of Decision Assurance
Research · No. 009 · Foundational

The Honest Limits of Decision Assurance

Abstract

Measuring what an AI system actually achieved is a better foundation than measuring what it did. It is not a solved problem. Some outcomes never arrive; some arrive late, partial, or subjective; reconciliation costs money and takes time; a feedback signal can itself be gamed; and separating a wrong decision from a right decision the world overturned is a hard, evidence-bounded judgment. This paper states those limits plainly and describes what a mature team does about each one. The argument is that a discipline honest about its boundaries is more trustworthy than one that pretends it has none, and that the disciplined response to an uncertain outcome is to say so, not to invent a cause.

Provy Research By the Provy Research team v2.1 12 min read Published July 2026 Updated July 2026 Download PDF ↓
Executive summary

Measuring what an AI system achieved is a better foundation for trust than measuring what it did, and it has real limits. Some outcomes never arrive, some come late or partial or subjective, the feedback signal itself can be gamed, coverage is bounded, and telling a bad decision from bad luck is hard. A discipline that names these limits is more trustworthy than one that hides them. The through-line: when the evidence is thin, the honest output is "we don't know yet," not an invented verdict.

Key takeaways
  1. Measuring outcomes is the right foundation, and it is bounded. Naming the bounds is what makes the measurement trustworthy.
  2. Some outcomes are unknowable, delayed, partial, or subjective, and some feedback signals can be gamed. Each has a disciplined response short of pretending certainty.
  3. You can only assure what you can observe the outcome of. Coverage is a real limit, and the honest number reports it rather than averaging it away.
  4. Separating "the agent was wrong" from "the agent was right and the world moved" is evidence-bounded. When the evidence is thin, the disciplined system says "insufficient evidence" instead of inventing a cause.

A discipline that only advertises its strengths is a sales pitch. This library argues that measuring what an AI system actually achieved is a sounder basis for trust than measuring what it did, and that argument is worth more if it comes with an honest account of where the measurement gets hard. Every serious practitioner already knows outcomes are messier than a clean pass/fail. Pretending otherwise does not build confidence; it spends it.1

So this paper is deliberately about the limits. Not as a disclaimer buried at the bottom of a page, but as the substance of it. Each section names a place where reconciling decisions against outcomes runs into something real, and each pairs the limit with what a mature team does in response. The through-line is a single stance: the disciplined answer to an uncertain outcome is to say the outcome is uncertain, not to manufacture a tidy verdict the evidence does not support.

Why a discipline should state its limits

Trust in any measurement rests on knowing what it does not cover. An auditor's opinion is credible partly because it is scoped: it says what was examined and what was not. A metric that claims to see everything is the one to distrust, because the failures it cannot see do not disappear, they just stop being reported. The value of naming a limit is that it converts a hidden blind spot into a known boundary, and a known boundary can be managed.

There is also a practical reason. Most of the limits below have a workable response, and you only get to the response by admitting the limit exists. A team that pretends reconciliation is instant never builds the interim judgment it needs while the outcome is pending. A team that pretends every decision has a clean outcome never learns to report coverage. Honesty here is not humility for its own sake; it is what lets the practice actually work.

Outcomes that are unknowable or never arrive

Some decisions have no observable outcome, ever. An agent recommends against pursuing a lead; the counterfactual, whether that lead would have closed, is unknowable because it was never worked. An agent chooses one of two valid phrasings for a contract clause and the deal proceeds; there is no world in which we see how the other phrasing would have fared. For a whole class of decisions, especially decisions not to act, the ground truth does not exist to be measured.2

What a mature team does. It does not fabricate an outcome to fill the cell. It marks the decision as one without an observable result and keeps it out of the numerators and denominators that imply one. Where a proxy outcome is genuinely informative, it is used and labeled as a proxy, never as the thing itself. The category of "no outcome available" is treated as a real, reportable state, not quietly rounded to success. An assurance number that silently counts unmeasurable decisions as wins is worse than no number.

The cost and latency of reconciliation

Comparing a decision to its real outcome is not free and it is not instant. The outcome lives in a different system and arrives on its own schedule. Pulling it back, matching it to the decision that caused it, and grading it takes work, and until that work is done the verdict is pending. There is a real gap between when an agent acts and when anyone can say whether it was right, and for some workflows that gap is long.

What a mature team does. It separates the two questions it can answer at two different times. At decision time it has an estimate, the system's own read on whether the work looks right. Later it has the reconciled truth. It reports them as distinct, resists treating the early estimate as the final answer, and is explicit that a pending outcome is pending, not passed. It also spends its reconciliation budget where it matters most, prioritizing the decisions whose outcomes carry the most consequence rather than trying to reconcile everything with equal urgency. The distinction between the two readings is developed as estimated versus reconciled.

A pending outcome is not a passing one. The discipline earns its trust by keeping "we don't know yet" visibly different from "it worked."

Outcomes that are delayed, partial, or subjective

Even when an outcome does arrive, it often arrives imperfect. It is delayed, so the record has to stay open long after the agent moved on. It is partial, resolving some of what a good result required and leaving the rest unknown. Or it is subjective, a matter of judgment rather than a fact that ties out.

Consider a procurement agent that approves a supplier contract. The clean part of the outcome arrives fast: the purchase order is issued, the goods are delivered. Whether the approval violated a sourcing policy that no one thought to check may not surface for two quarters, in an audit, and even then "was this the right supplier?" is partly a judgment call about risk appetite, not a number that resolves to true or false.3

What a mature team does. It refuses to collapse a partial or subjective outcome into a false binary. Grading a decision on many conditions rather than one verdict lets a result be "three of four conditions met, one still open" instead of a forced pass or fail. Subjective conditions are named as judgments and, where they matter, sent to a human rather than scored as if they were objective. The honest record shows what is settled, what is still open, and what is a matter of judgment, and does not pretend the three are the same kind of fact.

The risk of gamed or false reconciliation

A feedback signal is only as trustworthy as its source, and a source can be wrong or worked. The signal that an outcome "succeeded" might come from the same actor with an interest in that verdict. A ticket marked resolved may have been marked by the very agent whose work is under review. An outcome field can be stale, mis-keyed, or written by an upstream process that itself made a mistake. Reconciling against a corrupted signal launders a bad outcome into a clean one and does it with the full authority of a measurement.4

What a mature team does. It treats the outcome signal as evidence to be weighed, not gospel to be trusted, and it prefers independent sources over self-reported ones. Where a signal can only come from an interested party, that dependency is noted and the resulting verdict carries less weight. It watches for the tell of a gamed signal, an outcome metric improving while the thing it was meant to represent gets worse, which is the same specification-gaming pattern that shows up one level down in the agents themselves. The credibility of a reconciled outcome is a function of the credibility of its source, and a mature practice tracks the difference.

The coverage limit

This is the hard boundary of the whole discipline, so it is worth stating flatly: you can only assure what you can observe the outcome of. Decisions whose outcomes are unknowable, unobserved, or not yet connected sit outside the measured set. A team can have excellent numbers on the covered region and know nothing about the rest, and if it forgets that, the covered region's health gets read as the whole system's health.

What a mature team does. It makes coverage a first-class, reported number and holds it next to every quality figure. "Ninety percent of measured outcomes met the contract" is a different and weaker claim when the measured outcomes are forty percent of all decisions than when they are ninety-five. A confident score over thin coverage is treated as the provisional thing it is. The uncovered region is labeled unknown, not assumed fine, and expanding coverage is treated as the real work rather than tuning the score on the slice already in hand.5

Key idea

Coverage is the denominator honesty depends on. A quality number without a coverage number is a statement about an unnamed fraction of reality presented as a statement about all of it. The mature move is to always show both, and to let thin coverage weaken the claim rather than hide inside it.

The honest limit of attribution

Here is the hardest one, and the one where overreach does the most damage. When an outcome comes back wrong, the tempting next sentence is "the agent made a mistake." Sometimes that is true. Sometimes the agent made a defensible decision on the information it had and the world moved afterward, in a way no reasonable decision could have anticipated. A sales agent writes a well-judged follow-up, and the deal dies because the customer's budget was frozen the next morning by news the agent could not have known. The outcome is a loss. The decision was sound. Those are not the same finding, and treating every bad outcome as a bad decision is its own kind of failure.6

Separating "the agent was wrong" from "the agent was right and the world moved" is a genuinely hard, evidence-bounded judgment. It depends on what can be reconstructed about the decision, the information available at the time, and the events that followed, and often that evidence is incomplete. This library names the question, was the agent wrong, or did the world move?, as one the discipline must take seriously. It does not claim the question is always answerable, and it does not describe a mechanism that would resolve it in every case, because no honest one exists.7

What a mature team does. When the evidence supports a cause, it names the cause and holds the record that supports it. When the evidence is thin, it says insufficient evidence to attribute and stops, rather than reaching for the nearest plausible story. A blank where a cause should be is not a failure of the system; it is the system declining to assert something it cannot support. An attribution practice that always produces a confident cause is not thorough, it is unfalsifiable, and the difference between those is the whole value of the exercise. The concept and its boundary are treated directly in decision attribution.

On scope

This paper states that attribution is hard and evidence-bounded. It deliberately does not describe how any particular system attempts the separation, because the honest claim is about the limit, not a method that erases it. Where the evidence runs out, the correct output is "insufficient evidence," and any account that promises otherwise should be read with suspicion.

In practice

One idea, seen in a single kind of work. This illustrates the concept, not a product.

Security operations

A triage agent closes a low-priority alert as benign. Most of the time nothing follows, and "nothing followed" is not proof the call was right, only that no consequence surfaced. If a breach later traces back to that alert, the outcome finally arrives, months late. The honest record holds the decision as closed-benign with no confirmed outcome rather than counting it as a clean win, because the absence of a bad result is not the presence of a good one.

The honest stance, and why it is the trustworthy one

Put the limits together and a stance emerges. Some outcomes never arrive, so do not count them. Reconciliation is slow and costly, so keep the estimate and the reconciled truth distinct and spend the budget where it matters. Outcomes are often partial or subjective, so grade them as such and send judgments to humans. Feedback can be gamed, so weigh the source. Coverage is bounded, so report it. And attribution is hard, so say "insufficient evidence" when it is.

None of that weakens the case for measuring outcomes. It sharpens it. The alternative, measuring activity instead, does not escape a single one of these limits; it just hides them, by never asking the outcome question at all. A discipline that asks the question and is candid about where the answer is uncertain is on firmer ground than one that reports a confident green it never earned. Confidence that survives contact with its own limits is the only kind worth having, and it is the kind that Decision Assurance is trying to build.

The mature end state is not a system that always knows. It is a system whose "we don't know yet" and "we cannot attribute this" are as trustworthy as its verdicts, because they are used honestly and often. That is what separates assurance from a scoreboard, and it is why stating the limits out loud is not a weakness of the discipline but a condition of it.

Concepts in this paper

Notes

  1. The claim throughout the library is comparative, not absolute: measuring outcomes is a better foundation for trust than measuring activity. "Better" is not "perfect," and this paper is the accounting of the difference.
  2. Decisions not to act are the sharpest case: the counterfactual outcome is unobservable in principle, not merely unmeasured. No amount of instrumentation recovers a result that never happened.
  3. The procurement example is illustrative. The general pattern, an outcome whose fast part looks clean while its slow, judgment-laden part is still unresolved, recurs across compliance, credit, and long-cycle sales.
  4. A self-reported outcome shares an interest with the decision under review. Independence of the outcome signal is a property worth tracking, because a reconciled verdict is only as sound as the signal it reconciles against.
  5. Coverage functions as the denominator behind every quality claim. Reporting quality without coverage is reporting a rate without saying what it is a rate of.
  6. Treating every bad outcome as a bad decision punishes sound judgment for bad luck and teaches a system to avoid defensible risks. The error is symmetric with the failure of treating every good outcome as a good decision.
  7. Naming the question is not the same as claiming to answer it in all cases. The library takes the question seriously and is explicit that the honest output, when evidence is thin, is a declared absence of attribution.

References

  1. Provy Research. Decision Assurance: operating autonomous AI with confidence. Provy Knowledge Center, 2026.
  2. Provy Research. Why AI observability isn't enough. Provy Knowledge Center, 2026.
  3. Provy Research. AI Failure Modes. Provy Knowledge Center, 2026.
  4. Provy Research. Continuous Verification. Provy Knowledge Center, 2026.
  5. Provy Research. The Enterprise Confidence Gap. Provy Knowledge Center, 2026.

Change history

  1. v2.1 · July 2026 · Added the publication number, an executive summary, and an in-practice scenario.
  2. v2.0 · July 2026 · Rebuilt as part of the numbered research series.

How to cite

Provy Research. "The Honest Limits of Decision Assurance." Provy Knowledge Center, No. 009, v2.1, July 2026, provy.ai/knowledge/research/honest-limits.

See what an honest outcome picture looks like

Provy reports what reconciled, what is still pending, and how much of your real decisions it actually covers, so the number you trust is one that admits what it does not know.

Get a demo