Skip to content
LUNTA

Where the numbers come from

LUNTA · · 9 min read · Worked problem

Worked from published figures — not client artifacts

This piece works through the illustrative engagement published on our deliverables page — a service-desk deflection programme. Every figure in it is illustrative, including the ones derived here. We publish no client deliverables, and no anonymised ones either.

A pilot evaluation schedule is four rows and two signatures. It is also the most consequential document in an AI programme, because it fixes in advance what will count as evidence — and the reasoning that produced its numbers is almost never published with it. This is one worked through line by line, on the illustrative engagement we already publish.

Where the numbers come from — Schedule 2 — Pilot evaluation thresholds · illustrative
First-pass resolution on production tickets≥ 62%41% · measured in Diagnose
Answer accuracy, adjudicated sample (n = 500)≥ 95%—
P95 latency at operating volume≤ 4 s—
Cost per resolved case at volume≤ $0.40$3.10 · human-only

Only one number here was measured

The baseline was: 41% of tier-1 tickets resolved at first pass, counted during the diagnosis on the operation’s own ticket record, at a fully loaded variable cost of $3.10 for each case a person resolved. Everything else in the schedule is a decision taken against that measurement, before anyone knew what the system would do.

That is the property worth noticing. A schedule is not a forecast of how the build will perform. It is a statement of what would have to be true for the operation to have changed — written in the one week when writing it is cheap, because no result exists yet to make any particular number convenient.

Its four rows are not the same kind of number, though. One is arithmetic. One is judgement. One is a constraint on the architecture that the business case barely notices. And the number the case actually turns on is not in the schedule at all, for a reason worth defending. Saying which is which is most of the discipline.

The sample size is arithmetic

The accuracy row commits to 95% on an adjudicated sample of 500 cases. The sample is not there to estimate accuracy precisely — it is there to distinguish a system that clears the bar from one that misses it by an amount worth caring about, and those are different jobs with different costs.

An observed 95.0% on 500 adjudicated cases carries a 95% confidence interval of roughly ±1.9 points on the normal approximation: the true rate sits somewhere between about 93.1% and 96.9%, and a more conservative interval pushes the lower end to 92.7%. So this sample cannot settle a one-point argument, and a schedule written as though it could would be inviting one at the worst possible moment — after the result is known, with a launch date attached.

What it can settle is a shortfall of about two and a half points. Working the sample backwards from the difference worth detecting — at 80% power, against the 95% claim, on a one-sided test at the 5% level, using the normal approximation without a continuity correction — gives the size the schedule rounds to. The method is not decoration: state it and the table below can be reproduced by anyone who wants to argue with it; publish the sample size alone and they have to take it on trust, which is the posture this document exists to avoid.

Where the numbers come from — Sample size against the shortfall it can detect · illustrative
5 points — a true rate of 90%150Detected, with room to spare
3 points — a true rate of 92%383Detected
2.6 points — a true rate of 92.4%501One case beyond the schedule’s sample
1 point — a true rate of 94%3,118Out of reach, and out of adjudication budget

Note the third row, because it is the sort of thing that gets tidied away. Detecting a 2.6-point shortfall takes 501 cases, and the schedule commits to 500 — one short, because 500 is the number a human being writes down. That is a defensible rounding and an indefensible thing to leave unsaid: it is the difference between a sample size that was calculated and one that was chosen and then justified.

There is a sharper problem in the same row, and it only surfaces if you do the arithmetic. Sizing a sample from a test is committing to that test’s decision rule — and at 500 cases, a one-sided test of the 95% claim does not reject it until the measured rate falls below 93.4%. The schedule states something else: a hard bar, 95%, pass or miss. Those are two different rules and they disagree across a band 1.6 points wide. A pilot measuring 94.2% misses the schedule and clears the test it was sized by, which is exactly the argument nobody wants to be having in the week the result lands.

The hard bar has its own edge, too. A system whose true accuracy is precisely 95% — the standard the schedule sets — clears a strict reading of that bar on 500 cases only about 55% of the time, because a measurement is a sample and samples fall on both sides of the thing they measure. Neither rule is wrong. What is wrong is a schedule that implies one, is sized by the other, and names neither: it states the measure, the threshold, the baseline, and the sample, and it does not state which rule governs when they disagree. Four of the five. The fifth is the one that gets argued about.

The binding constraint is not the statistics, though; it is the adjudication. Every one of those cases is read and judged by a person who knows what a correct answer looks like. At four minutes a case — the assumption the schedule was costed on — 500 adjudications is about 33 hours of a senior handler’s time, and 3,118 would be over five working weeks of it, which no operation is going to fund in order to settle one point.

So the honest sequence is: decide the shortfall worth detecting, compute the sample, price the adjudication, and if the price is more than the answer is worth, change the measure or the threshold rather than quietly hoping nobody asks what the sample can support. The schedule names who adjudicates for the same reason it names the sample size: both are part of what the number means.

The resolution threshold is a judgement, and it is not the breakeven

The value model behind this engagement reconstructs from the figures already published in the closing statement, and it is worth doing explicitly, because the reconstruction is what makes the next two sections checkable rather than assertive.

Published at close
First-pass resolution 64% · cost per resolved case $0.38 · annualised net benefit $1.9m, after $0.4m of annual platform and model run cost.
Unit-cost delta
$3.10 − $0.38 = $2.72 saved on each case resolved without a person.
Cases the delta applies to
$1.9m net + $0.4m run cost = $2.3m gross ÷ $2.72 ≈ 845,000 cases a year.
Ticket volume implied
845,000 ÷ 64% ≈ 1.32 million tier-1 tickets a year. The closing statement does not name a volume; this is the volume its own arithmetic commits to.
What the model assumes
That every case resolved at first pass after the system arrives is a case a person would otherwise have handled, credited at the average human cost. Two effects pull against each other inside that average and the published figures do not settle which is larger: deflection takes the routine cases first, which are cheaper than the average and so overstate the saving, while the cases that used to escalate cost more than the average and so understate it. We are not going to claim the model is conservative when what we can actually say is that it is unadjusted.

With that model in hand, the breakeven is computable: the first-pass resolution rate at which the programme merely pays for its own run cost is $0.4m ÷ (1.32 million × $2.72), or 11.1%. Just over eleven tickets in a hundred, and the programme washes its face.

Which is exactly why 62% is not derived from the money. A threshold set at breakeven licenses a system that pays for itself and changes nothing: a second path to operate, monitor, secure, and explain, in exchange for a saving the operation could not find on its own accounts. And note where that line falls — far below the 41% the operation already manages without any of this. The money does not constrain the threshold at all, which is precisely why the threshold has to come from somewhere else and be seen to.

62% is one and a half times the measured baseline, rounded up to the next whole point. That is a judgement — the point at which the operation is doing something materially different from what it does today — and stating it as a rule rather than a number is what makes it arguable in the week it costs nothing to argue about. A sponsor who thinks half again is too easy, or too hard, can say so while the number is still a sentence.

The cost ceiling polices the architecture, not the case

Now the row that looks like money and is not. Moving one variable at a time through the model above shows what each measure is actually worth to the case.

Where the numbers come from — What moves the case, and by how much · illustrative
As published at close — 64% and $0.38$1.90m—
Cost per resolved case at its ceiling — $0.40$1.88m−0.9%
First-pass resolution at its threshold — 62%$1.83m−3.8%
Both schedule measures at their limits$1.81m−4.6%
Annual run cost comes in at double — $0.8m$1.50m−21.1%
Ticket volume a fifth below the diagnosis figure$1.44m−24.2%

The measure the schedule polices most tightly turns out to be the one the case is least sensitive to. Running at the $0.40 ceiling instead of the measured $0.38 costs under one per cent of the benefit. Both schedule measures at their limits together cost 4.6%. Meanwhile a fifth off the ticket volume costs 24% — more than five times as much as both thresholds being scraped rather than cleared.

That is not an argument for dropping the ceiling. It is an argument for knowing what the ceiling is for. Unit cost is an architecture decision made early and discovered late: which tier serves which class of request, how much retrieved context each call carries, whether the prompt prefix is stable enough to cache, whether the routine majority can run on a smaller model and escalate only on low confidence. A number in the schedule is what forces those decisions during the build rather than during the rollout, and it has to be tight enough to still bind at ten times the volume and on a model generation that does not exist yet. A ceiling set where the business case stops caring is not a ceiling. It is a rounding error with a signature under it.

What the schedule does not contain

Volume. The single input the case is most sensitive to appears nowhere in the schedule, and that is correct rather than an oversight: a pilot cannot prove how many tickets the operation receives. That number is measured during the diagnosis, on the operation’s own record, and it is one of the reasons the diagnosis is a separate piece of work with its own artifacts.

So the two documents divide the risk between them. The diagnosis carries the question of whether the value is there at all — volume, baseline, unit cost, data reality. The schedule carries the question of whether the system earns the right to go after it. Signing one and skipping the other produces the two most common failures in enterprise AI: a well-evidenced system nobody needed, or a genuine opportunity attacked with a build that was never judged.

The discipline is saying which kind of number each row is

Four rows, three kinds of number. The sample size is arithmetic, and it can be recomputed by anyone who disagrees with it. The resolution threshold is a judgement expressed as a rule, and it should be argued about before the pilot rather than after it. The cost ceiling is a constraint on the design whose value to the business case is close to zero and whose value to the system is most of why it is still running in three years. Latency is the fourth, and it belongs to the operation rather than to any of the above: a correct answer arriving after the handler has already picked up the phone did not resolve anything.

A schedule whose rows do not say which they are invites the wrong argument at the wrong moment — the judgement gets defended with arithmetic, the arithmetic gets defended with conviction, and the row that was a design constraint gets traded away by someone reasonably observing that it barely moves the number. Labelling them costs a sentence each.

None of this is difficult, and none of it is proprietary. It is simply written down before the result exists, which is the only moment any of it is cheap.

Read next

The honest no

Evaluation gates exist to kill workstreams. If yours can’t, they aren’t gates — they’re theatre.