Oct 9, 20268 min read/2026/10/09/benchmarking-a-decider-accuracy-and-calibration/

Benchmarking a Decider: Accuracy Is Not Calibration

A vendor tells me their decision model is 94% accurate. My first reaction used to be "great, let's use it." Now it's "94% accurate at what, on whose examples, and what does it do with the other 6%?"

I've been building "System One" deciders for a while now — you hand a model some state, a question, and a fixed list of options, and it hands back one of those options with a probability. I wrote about what Jev is, about putting Jev, a local model, and a fake behind one IDecider, about the alternatives that actually replace it, and the System 1 / System 2 framing. What I never wrote down is the part that decides whether you can ship one: how do you score it on your own problem?

The restaurant critic

Picture a food critic who has reviewed a hundred restaurants and never once eaten at yours. That review is careful, methodology-driven, and useless for the only question you have: is this place good for me tonight?

A vendor's benchmark number is that review. It was earned on their test set, chosen and labelled by them, with their definition of the classes. Your tickets are your restaurant. The fix isn't to distrust the number — it's to run your own review: twenty or thirty examples out of your actual traffic, labelled by you.

Accuracy alone cannot tell you what to do with the 6%

A decider is not a classifier. Its output is a value and a confidence, and if you only measure accuracy you have thrown away the half you were going to build your production logic on.

Think of a forecaster who says "rain tomorrow" every day. Where it rains 40% of days, that forecaster is 40% accurate and useless, because it never tells you how sure it is. A good forecaster isn't the one who is right most often; it's the one whose 70% means 70%.

That property is calibration, and it's what makes abstainBelow work. Say 0.9, decide; say 0.55, escalate to a human or a bigger model. That logic is only sound if the numbers mean something.

So the benchmark reports four things per engine:

Metric Question it answers
accuracy how often is the top choice right?
mean confidence how sure does it claim to be?
gap (confidence − accuracy) is it over- or under-confident overall?
ECE is it honest per confidence band, not just on average?

ECE, without the ceremony

Expected Calibration Error is not exotic. Bin the predictions by the confidence they claimed, compare claimed against actual in each bin, then average the gaps weighted by bin size. For a bin \(B_m\):

\(\text{ECE} = \sum_{m=1}^{M} \frac{|B_m|}{n} \Big| \text{acc}(B_m) - \text{conf}(B_m) \Big|\)

Zero is perfect — "when it says 90%, it is right 90% of the time." The metric comes from Guo, Pleiss, Sun and Weinberger's 2017 paper on calibrating modern neural networks; that paper used image classifiers, but the arithmetic doesn't care what produced the probability.

One warning I learned the hard way: with a few dozen examples, bin choice matters. Ten equal-width bins over 30 predictions leaves most empty and the number jumps around. So the harness computes equal-width and equal-mass binning (equal-mass puts the same count in each bin, so an empty bin can't hide a problem). Here's the equal-width version:

// ECE: bin the predictions by the confidence they claimed, then average
// |accuracy - confidence| per bin, weighted by how many predictions landed in it.
static double Ece(IReadOnlyList<Score> scores, int bins)
{
    double ece = 0;
    for (int b = 0; b < bins; b++)
    {
        double lo = (double)b / bins, hi = (double)(b + 1) / bins;
        var inBin = scores.Where(s => s.Confidence > lo && s.Confidence <= hi).ToList();
        if (inBin.Count == 0) continue;
        double acc = inBin.Count(s => s.Correct) / (double)inBin.Count;
        double conf = inBin.Average(s => s.Confidence);
        ece += inBin.Count / (double)scores.Count * Math.Abs(acc - conf);
    }
    return ece;
}

That's the whole metric. No library, no framework.

The harness: your own labelled set

The expensive part of benchmarking is never the code — it's the labels. So make the set small enough that you'll finish it. Mine is support-ticket routing: three teams, and Which team should handle this?

record Case(string Ticket, string Expected, string Difficulty);
record Score(bool Correct, double Confidence, string Difficulty);

var cases = new List<Case>
{
    new("I was charged twice for the same month and the invoice shows two line items.", "billing",   "clear"),
    new("The API returns 401 on every request even though my key is correct.",           "technical", "clear"),
    new("Can I get a demo for my team of twelve next Tuesday?",                          "sales",     "clear"),
    new("Rate limit is 100/min but the docs say 600/min; which is right?",               "technical", "ambiguous"),
    new("It does not work.",                                                             "technical", "ambiguous"),
    // ...
};

The Difficulty field is the second thing I got wrong. I used to label everything "clear" and then wonder why my scores looked good and production looked bad. Half of real traffic is the vague tickets — "it does not work", "can someone call me", a message that is genuinely two problems at once. Those are the cases where a decider either admits uncertainty or invents confidence, so label them and report them separately. Averaging easy cases with hard ones hides exactly the behaviour you're trying to buy.

The scoring loop is deliberately boring: call the interface, compare the top value to the label, keep the confidence.

foreach (var c in cases)
{
    var d = await decider.ChooseAsync(c.Ticket, question, teams, abstainBelow: 0.0);
    bool correct = d.Value == c.Expected;
    scores.Add(new Score(correct, d.Confidence, c.Difficulty));
}

Because everything sits behind one IDecider, the same loop scores Jev, open-jev, a llama.cpp model under a GBNF grammar, strict json_schema over Azure/Ollama/vLLM, or a FakeDecider. Swapping the engine is one line at the top of the file. That's the payoff of the interface — not elegance, but comparability. You cannot compare two engines measured by two different scripts.

What the output looks like

Here's a real run. To keep it reproducible without a key or a GPU, it points at a stand-in that behaves like a vendor pitch: always picks billing, and reports a confidence that rises with how much text it was given.

all cases  (n=22)
  accuracy                      31.8%
  mean confidence               91.8%
  gap (confidence - accuracy)   60.0%   OVERCONFIDENT
  ECE  (equal-width, 10 bins)  60.00%
  ECE  (equal-mass,   5 bins)  60.00%

  reliability (equal-width bins)
    confidence band    n   claimed   actual
     80% -  90%       4       87%      25%
     90% - 100%      17       94%      35%

Read the reliability table, not the headline. The decider claimed 94% on 17 tickets and got 35%. Nobody looking only at accuracy would see that shape — and the shape is the shipping decision. If your production code routes whenever confidence exceeds 0.8, this decider hands 21 of 22 tickets to an automated path while being wrong most of the time.

Split the same tickets by difficulty and it gets worse where it matters. On the 16 clear cases: 37.5% accuracy against 93.4% mean confidence, ECE 55.87%. On the six ambiguous ones — "it does not work" — 16.7% accuracy against 87.7% confidence, ECE 71%. The vague tickets are where it's worst and where it is most convinced, which is the failure mode that makes a decider dangerous rather than merely inaccurate.

Three honest caveats

First, my stand-in is synthetic. Its accuracy is low and its confidence comes from a formula, which is why the gap is a cartoonish 60 points. The harness, metrics and equal-mass ECE are real, though, and the stand-in needs no key — point it at a real engine with the first CLI argument and the report is identical. For a headline you can quote, get it from a real engine on your own data.

Second, the gap and ECE agreeing at 60% in that first table is a coincidence of the data's shape, not a redundancy. With confidence spread across many bands and errors unevenly distributed, the two diverge — and ECE is the one that catches a decider calibrated on average but badly wrong in one band.

Third, this benchmark does not measure abstention or cost. It runs with abstainBelow: 0.0 so every case produces a decision. The next post I need sweeps the threshold and asks what the accuracy is on the tickets the decider chose to answer, how many it handed off, and what escalation costs per thousand.

Stop trusting the number, run the review

None of this is difficult, which is the real argument for doing it. The harness lives at github.com/egarim/systemone-deciders under samples/Sample.Bench — a case list, a scoring loop, Ece, EceEqualMass, and a reliability table. dotnet run --project samples/Sample.Bench needs no API key and no GPU. Twenty labelled examples from your own traffic and you can stop arguing about which model is better and start measuring which one is better for your tickets.

A vendor's accuracy number is a review of somebody else's restaurant. Go eat your own dinner.

If you're scoring deciders on real traffic — or you've found a confidence band where yours quietly falls apart — tell me through the about page.