S01E03: Confidence, Calibration & Composite Scoring

  • Level L2 (intermediate)
  • prerequisite: S01E01|~20 min|verified 2026-09-24
  • calibrationjevcomposite-scoringthresholdsllm-as-judge

Every Jev answer arrives with a number attached, and the entire product pitch is that you can automate against that number. This lesson covers what the number promises, how to turn probabilities into decisions with thresholds, how to rank with composite scores, and how to check the promise with measurements instead of trust.

What the numbers promise

Jev's real product claim is not speed. It is calibration: higher confidence should track higher accuracy, and the model is trained to make that true.

🏷️ New termCalibration· glossary

A judge is calibrated when its confidence numbers track actual correctness: among the answers it scores at 0.8, roughly 80 percent are actually right. Calibration is what makes confidence usable as a probability, a number you can build thresholds on.

Why would a model be calibrated at all? Because it was trained to be. Jev's training objective is RLCD (Reinforcement Learning for Calibrated Decisions), which rewards honest probabilities instead of confident-sounding ones (vendor's framing, launch post). It targets a known mainstream defect: RLHF demonstrably hurts calibration (the GPT-4 technical report's reliability plots show it), and a judge that is confidently wrong is worse than one that hedges honestly. The idea that a model can be trained to know what it does not know predates Jev: Kadavath et al. established it in 2022 (further reading below).

🏷️ New termRLCD· glossary

Reinforcement Learning for Calibrated Decisions, the training objective that rewards honest probabilities instead of confident-sounding ones.

Here is why this matters more than accuracy. Consider a judge that automates a decision: "act when confidence ≥ 0.9". If the model is calibrated, that line means "act when it is right at least 90 percent of the time", and you can tune it with data. If the model is miscalibrated, its 0.9 might be right only 60 percent of the time, and your threshold enforces a policy that does not exist. High accuracy with broken calibration still breaks threshold-based automation, because the number your code acts on stops meaning what you think it means. That is why a miscalibrated judge is worse than its accuracy suggests.

💡 Key insight

Calibration is a measurable property. Measure it instead of trusting it: a model that is accurate but miscalibrated fails the judge use case even though its verdicts look right. Section 5 shows the two standard instruments.

Reading a Noul against thresholds

A noul answer is a probability with no separate confidence field: the value itself is the uncertainty signal. Reading it against a threshold policy turns it into a verdict:

Noul valueWhat it meansReasonable action
≥ 0.95Confidently trueCount as met, move on
0.70 – 0.95Probably truemet under a precision-biased threshold of 0.7
0.30 – 0.70Genuinely uncertainTreat as the abstention band; route to review
0.05 – 0.30Probably falseunmet
≤ 0.05Confidently falseunmet, no review needed
🏷️ New termAbstention band· glossary

The Noul range 0.30–0.70, where the model is torn between the two answers. The fraction of answers landing there is worth reporting: it is the share of work a human still has to look at.

The split the table enforces is the important design move: the model supplies probabilities, and the threshold policy is yours. The same assessment can feed a strict policy (act only at ≥ 0.7, encode a "when in doubt, fail" bias) or a lenient one (act at ≥ 0.5, accept more noise). Where the cutoff sits is a product decision you own in code, not prompt wording.

From probabilities to decisions: threshold routing

One battery of Noul questions plus one severity score is enough to route real traffic, and the vendor's guardrails cookbook works the full example (all figures below are from the vendor cookbook, verified 2026-09-23; numbers produced by jev-1.12 on 2026-08-15). The shape: screen a message with a battery of hazard nouls (jailbreak, harmful request, medical advice, self-harm) and one severity score. Two thresholds do the routing:

  • at or above the action threshold, the hazard triggers its configured action;
  • at or above the lower review threshold, the message goes to a human;
  • below both, it passes unless another hazard fires.

The severity score has a threshold of its own and can escalate a review into a block. A "policy" is just those numbers under a name, which turns the trade-off into something the application picks rather than inherits:

# Same TypeSafe result: jailbreak=0.74, severity=0.51

strict       review >= 0.35  action >= 0.70  ->  block
permissive   review >= 0.35  action >= 0.85  ->  review
ONE ASSESSMENT jailbreak = 0.74 severity = 0.51 cached; the model is not called again strict policy review ≥ 0.35 action ≥ 0.70 permissive policy review ≥ 0.35 action ≥ 0.85 block refuse the turn review a human looks the probabilities do not move; the policy decides how much evidence to demand

The cookbook's own numbers show why the routing layers earn their keep. Under the strict policy, one message asking about a melatonin dose landed at medical_advice=0.55 and went to a human (a dosage question mild enough to hand to a person), while an 800 mg ibuprofen dosage reply crossed the severity block line at 2.02 and became a block. A message dressed up as a medical accommodation scored jailbreak=0.74: the strict policy blocked it, the permissive one would only review it. Same probabilities, different decisions. The assessment is reusable; only the thresholds change.

💡 Key insight

TypeSafe supplies the assessment; your application owns the decision. The four-way route (pass, review, block, support) is what a plain block cannot do: it is the difference between helping someone and hanging up on them. Both the hazard battery and the thresholds are yours to edit.

Uncertainty gates: when the answer is not settled

The thresholds above assume a Noul, whose single probability is both the evidence and the uncertainty statement — the abstention band (0.30–0.70) falls out of it for free. Choice and Score carry two dials instead (S01E01), so "in the band" becomes a predicate — a small OR of conditions, any one of which routes to review:

// Choice: margin = winner − runner-up; both gates needed
inBand = confidence < minConfidence  ||  margin < minMargin

// Score: relative to the decision, not intrinsic
inBand = confidence < minConfidence      // hedged position
      || tornDistribution               // mass split across adjacent levels
      || nearThreshold(threshold)       // straddling the tripwire
🏷️ New termmargin· glossary

The winner's probability minus the runner-up's. The coin-flip detector for a Choice: the answer always looks like an answer, but a 0.51/0.49 winner is arbitrary between ties. Margin is what tells a decisive 0.85/0.15 from a torn 0.51/0.49 with the same winner label.

The two Choice gates catch different shapes of unsettled: 0.9/0.05/0.05 has a clear winner but low confidence; 0.51/0.49 may carry decent confidence on the winner while the top two sit in a dead heat. Margin catches the second, confidence the first — you need both. For a Score, "in the band" is relative to the decision, not intrinsic: 1.4 is unsettled only if your threshold is nearby, and a Score at 1.9 against a 2.0 block line with torn mass is not a pass — it is a maybe, and maybes go to review.

💡 Probabilities choose the action; confidence chooses whether to automate it

Because confidence is calibrated, low confidence is a prediction of error, not noise: when the model says 0.55, it is wrong about 45% of the time. So the working rule is two dials in two roles — the probabilities (or the Noul value) decide what to do; the confidence (and the margin) decide whether that action is safe to automate. An unsettled answer has three exits, in order of preference: route to review, take the safe default (the cheaper-to-fail option), or rewrite the question. A recurring tie or a wide band is not the model's indecision — it is the menu, the levels, or the criteria wording, and the tie rate is the metric that measures your schema. Never "interpret harder": act, review, or rewrite.

Composite scoring: weights live in your code

Routing handles one binary outcome at a time. Ranking several items needs more, and the composite scoring pattern (vendor pattern doc, verified 2026-09-23) is the designed shape for it: break the judgment into independent dimensions, score each separately, and combine with weights you control in code.

The vendor's worked example ranks job candidates on four score questions in one request: Python depth, team leadership, system design, generalist range. Each dimension has written levels from "no experience mentioned" to "deep expertise". Then, entirely in code:

py      = response.answers["python_depth"].score / 4
lead    = response.answers["team_leadership"].score / 4
arch    = response.answers["system_design"].score / 4
general = response.answers["generalist"].score / 4

# Senior IC
ic_score = (0.40 * py) + (0.10 * lead) + (0.40 * arch) + (0.10 * general)

# Engineering Manager
em_score = (0.15 * py) + (0.40 * lead) + (0.20 * arch) + (0.25 * general)

Three properties make the pattern work:

  • Normalization is yours. Each score is divided by 4 (the highest level index) to land in 0–1 before weighting.
  • Weights are configuration, not prompts. The same four answers serve two different roles (senior IC, engineering manager) under two weight sets. Changing the ranking logic never touches the model.
  • The arithmetic is visible. You can see exactly how a final score was produced, and adjust weights when the top-ranked candidates do not match expectations.

This is also where the independence property from Lesson 1 pays off: the four dimensions are scored in parallel in one request, none influenced by another's answer, so a resume that games the leadership dimension does not move the Python score.

🔧 What composite scoring does not give you

Composition multiplies coverage; it does not create independence between models. Ten questions share one model's blind spots, so the weights combine correlated errors as readily as signals. If the ranking matters, hold back a labeled sample and check the composite against it before you trust the ordering (Lesson 1's "correlated mistakes" edge, restated where it bites).

Measuring calibration instead of trusting it

Two standard instruments turn "should be calibrated" into a number you can act on:

🏷️ New termReliability curve· glossary

Predicted probability buckets against observed accuracy. Bucket answers by stated confidence, measure how often each bucket was right, and plot the two against each other. The diagonal is perfect calibration.

🏷️ New termBrier score· glossary

A proper scoring rule: the mean squared difference between stated probability and outcome. It punishes confident wrong answers hardest, so a well-calibrated model scores lower than a lucky one.

The reliability curve shows you where the model is off (a bucket that says 0.9 but is right 70 percent of the time); the Brier score compresses the whole curve into one comparable number. Measure both against human labels on your own inputs, not the vendor's benchmarks: calibration claims are trained behavior, and your input mix is the distribution that matters.

One more instrument belongs in the set when the task is agreement with human judgment: κ (Cohen's kappa), agreement between two raters beyond what chance would produce. It is the standard metric for "does the model's verdict match a careful human's verdict" on binary met/unmet decisions, and it is what a judge's calibration monitoring should track over time.

🏷️ New termκ (Cohen's kappa)· glossary

Rater agreement beyond chance. Chance-adjusted, so it does not reward agreement that mere class balance would produce.

🔧 The operating loop

Set the threshold policy from labeled examples → run the judge → track κ and the reliability curve per run → when your input mix changes, expect the numbers to move and re-check before acting on them. Calibration is trained behavior plus your distribution; both halves need watching.

What the promise does not cover

Three edges keep the calibration claim honest:

The abstention band is a design decision

Everything between 0.30 and 0.70 is genuinely uncertain, and automation has to decide what happens there. The guardrails pattern routes that band to human review; a forced-choice schema with no other option instead forces an answer even when the true answer is missing from the menu. The schema author carries that burden (Lesson 1), and the abstention-band fraction is the metric that tells you how much of your traffic is quietly waiting for a human.

Calibration claims are vendor-measured

The RLCD story, the latency claims, and the headline eval numbers all come from the vendor's own workflows. Treat them as ceilings, and measure on your own labels. A vendor that trains for calibration has an incentive to be right about it, but your distribution is not their benchmark.

🏷️ New termOOD drift· glossary

Out-of-Distribution drift: inputs and policies shift over time, and probabilities that were reliable last month can move without notice. The countermeasure is the monitoring loop from section 5: κ and reliability curves, tracked on every run.

Confidence is per answer, not per request

Each question in a request carries its own probabilities. A batch of five answers is five separate calibration events, not one confident verdict; report the band or the minimum confidence across the battery, not an average that hides the one answer the model was torn on.

Knowledge check

📖 This knowledge check is an interactive quiz (5 questions). Open this page in a browser to take it; reader mode shows only the prose.

Further reading

🧭 Practice

Your turn. Take one binary decision you already automate (or would like to) and write its threshold policy down: what noul value counts as met, what band routes to review, and what the action threshold encodes about your error costs. Then open the TypeSafe Playground, run one real input through it, and check whether your thresholds would have routed the input the way you intended.

If an AI assistant is part of your workflow, Copy page in the header menu hands it this whole lesson as markdown, citation included. Ask it to quiz you on the knowledge check, or to explain a section a different way. The glossary is the canonical vocabulary when terms blur.

See also