Knowledge Center/ Research/ How Provy Checks the Work
Research · Field guide · Follow one run

How Provy Checks the Work

In one line

For every run, Provy keeps one record that starts as a guess about how the run went, then gets the real result stamped onto it once it comes in. Reconciliation is that second step: lining the guess up against the result, so a confident guess becomes a fact. The rest of this guide follows one support ticket through Provy so you can see it happen.

Provy Research By the Provy team v2.0 · draft 5 min read Unpublished draft Download PDF ↓
PROVY'S GUESS LINE THEM UP TRUST The run what the agent did, step by step How good it looked A pass / fail call, with a confidence The contract, graded on the claim GUESS vs REAL right or wrong, part by part The real result from your own records Trust score how much you can rely on this fleet
Provy makes a guess from the run, then checks it against the real result.
The short version

Before any outcome, Provy already grades every run from its trace: a quality read that guesses pass or fail, and the contract checked against what the agent claimed. That is the estimate, and it is labeled a guess. Later the real result comes back from your own records, and Provy grades the contract a second time against reality, then lines the two grades up, one condition at a time. A condition that looked met on the agent's claim but missed on reality is a divergence, and it is what pulls the trust score down.

Four things to hold onto
  1. Before the outcome, Provy grades the run from its trace: a quality guess, and the contract checked against what the agent claimed. That is the estimate.
  2. The real result comes from your records. You connect it; Provy does not read your systems.
  3. Reconciling grades the contract a second time against reality, then lines the two grades up part by part. A run passes only if every part held on reality.
  4. The scary case is a run that looked great and was wrong. Only reconciling catches it, and it is what pulls the trust score down.

Meet the run

A customer writes in asking for a refund. Your support agent reads the message, checks the account, and replies: refund approved, forty dollars back to the card. The reply is polite, clear, and fast. The customer says thanks. On every screen you have, this looks like a good run.

We will follow this one ticket the whole way through. Keep the forty-dollar refund in mind.

Step 1 — Provy grades the run up front

The moment the run finishes, the result is not in yet. Nobody knows if the refund was right. But Provy does not wait to have an opinion. From the run alone, it grades what it can, in three parts, all read straight from the trace. Together they are the estimate: how the run looks before reality weighs in.

How good it looked
Provy scores the quality of the run: was it correct, was it grounded, did it follow the steps. Here it scores high. The reply reads well.
A pass / fail call, with a confidence
From that quality, Provy guesses pass, and it is fairly sure. Call it 82% confident. The better it looked, the more sure the guess.
The contract, graded on the agent's claim
Your contract lists what a good ticket looks like. Provy checks each part against what the agent reported: refunded $40 and treated it as within policy, resolved the ticket, no duplicate. On the agent's own numbers, all three look met. Looked: 3 of 3. A clean-looking pass.
This is the estimate: how the run looks from the inside, before reality weighs in. Notice the contract already has a grade here, 3 of 3, built entirely on the agent's own claim. Hold that number.

Step 2 — the guess waits

At this point nothing is confirmed. The forty dollars has not been checked against anything real. On the Ledger this run sits as at risk: predicted good, but not proven. It is not counted as money earned and not as money lost. It waits for the real result. If a result never shows up, Provy eventually files it as unresolved, which is not a pass and not a fail, just an item with no ground truth. Provy does not invent one.

Step 3 — the real result comes in

A few days later, your finance system settles the refund. Your records show the plan the customer was on allowed a twenty-five dollar refund, not forty. The agent gave away fifteen dollars it should not have. You send that result back to Provy, tagged to the same ticket. Provy does not reach into your finance system; you push the record, and the ticket id is what lets Provy find the exact guess this result belongs to.

Step 4 — Provy lines the two grades up

Now the real result is stamped onto the same record, next to the estimate. Provy grades the contract a second time, this time against reality, and puts the two grades side by side. Up front, on the agent's claim, it looked like 3 of 3. Against the real refund record, one condition flips.

THE CONTRACT, GRADED TWICE · one ticket Contract part Looked (agent's claim) Really (reconciled) Refund within policy $40, ok cap was $25 FLIPPED Resolved, no callback resolved no callback No duplicate refund none none Looked 3 of 3Really 2 of 3 · the one that flipped cost $15
The same three parts, graded on the claim and again on reality. The gap is the flipped row.

That flipped row is the whole point. "Refund within policy" looked met on the agent's own numbers and was not, once the real cap came back. A run passes only when every part holds on reality, so two of three makes the ticket a miss, priced at its real fifteen-dollar cost. On the Ledger it moves out of at-risk and into lost. And the up-front pass/fail guess that said pass? Reality said fail, so that guess diverged too. Reconciliation is just this: grade on the claim, grade on reality, and see what moved.

Four ways it can land

Every reconciled run falls into one of four boxes: Provy guessed pass or fail, and reality passed or failed. This is the whole game, and it is worth seeing at once.

Really PASSED Really FAILED Provy guessed PASS Provy guessed FAIL Right A clean win. The run looked good and delivered. Confident, but wrong The $15 refund. Looked perfect, passed every check, still lost money. Only reconciling catches this. Cautious Provy flagged it as weak, but it turned out fine. Right Provy called the miss. It looked weak and it did miss.
The top-right box is the one that costs you, and the one only Provy catches.

Two of these boxes are Provy being right: it liked a run that delivered, or it doubted a run that missed. The bottom-left is Provy being cautious for nothing. The one that matters is the top-right: the run looked great, sailed through every check, and still lost money. Our refund ticket lives there. That is the box no quality score and no eval can see, because everything about the run looked fine. Only lining it up against the real result finds it. It is also the box behind the Command Center's "confident, and wrong" number.

The one that matters

So Provy does not stop at "this run missed." For a run that looked fine and was wrong, it pinpoints the step behind it. In our ticket, the miss traces back to the policy lookup, which handed the agent the wrong refund cap. Provy names that step, so you get a specific thing to fix instead of a number that dropped. When a run holds no clear cause, Provy says so plainly and calls it a blind spot to instrument, rather than inventing a reason.

How it becomes trust

One ticket is a story. Thousands of them are a trust score. Reconciliation is the engine under that score in three ways, all in plain view on the product.

It weights toward reality. Early on, with few real results in, the score leans on how runs look. As more results reconcile, it leans on what actually happened. The score earns its way toward the truth instead of claiming it on day one.

It watches the guessing. If Provy's guesses keep running ahead of reality, the fleet is overconfident, and the score comes down. That is the "how it looked versus what proved" bar on the Command Center.

It shows the gap. On the Ledger the headline reads "Looked X% good. Really Y%." The first number is how the runs looked. The second is how many actually held up once reconciled. The distance between them is the guessing running ahead of the truth, and closing it is the job.

Why a green "Trusted" waits

A green badge says you can rely on this fleet. Provy will not show it on guesses alone. Until enough real results have come back to stand behind the claim, the score reads Estimated or Calibrating. It goes green only once reconciliation has earned the right to say so.

The same pattern, other work

Our example was a refund, but the shape holds anywhere an agent's work has a real result. Something looks met on the agent's own claim, then misses once reality comes back. That flip is the miss, and it is the part no quality check sees.

Trading
Looked: the trade closed in profit.  Really: it breached the drawdown limit and lost on settlement.
Claims
Looked: claim approved, coverage matched.  Really: paid at the wrong rate for that plan version.
Sales
Looked: lead qualified and routed to the right rep.  Really: a duplicate account that never converts.

Same run, same estimate, same reconcile. Only the words on the contract change.

What Provy will not do

Reconciliation is defined as much by its restraint. It never invents a result that has not arrived; a run with no reported outcome waits as unresolved, not a pass or a fail. It never reads your systems; you connect the real result, tagged to the item. And it never blames a step it cannot see in the run; an unexplained miss is called a blind spot, not a guess. A number you can stand behind is worth more than a number that merely looks good.

Notes & references

For the one-page picture of the whole product, see How Provy Works. For the short definitions, see the concept pages Reconciliation and Estimated vs Reconciled. The idea of a run that looks fine and is not is drawn out in the insight Confident and Wrong.

Version history

  1. v2.0 · draft · July 2026 · Rewritten in plain language around one worked example. Not yet published.
  2. v1.0 · draft · July 2026 · First draft of the reconciliation walk-through.

How to cite

Provy Research. "How Provy Checks the Work." Provy Knowledge Center, Field guide, v2.0 draft, July 2026, provy.ai/knowledge/research/how-provy-checks-the-work.

Watch one run go from guess to fact

Send a run, connect its result, and see Provy's guess line up against reality part by part, then move the trust score.

Get a demo