---
title: "Confidence, Calibration & Composite Scoring"
description: "What Jev's confidence numbers promise, threshold routing, uncertainty gates for torn answers (margin, low confidence, torn distributions), composite scoring with weights in code, and measuring calibration instead of trusting it."
license: "© 2026 Peripatos — free to read; quoting with credit welcome, reuse by permission. Unofficial educational resource; Jev and TypeSafe are trademarks of TypeSafe AI."
dateModified: 2026-09-23
source: "https://peripatos.dev/courses/jev-fundamentals/s01e03-jev-confidence-calibration-composite/"
---

*Level L2 (intermediate) — prerequisite: S01E01 · ~20 min · verified 2026-09-24 — Tags: calibration, jev, composite-scoring, thresholds, llm-as-judge*

Every Jev answer arrives with a number attached, and the entire product pitch is that you can automate against that number. This lesson covers what the number promises, how to turn probabilities into decisions with thresholds, how to rank with composite scores, and how to check the promise with measurements instead of trust.

## In this lesson

1. [ What the numbers promise](#promise)
2. [ Reading a Noul against thresholds](#reading)
3. [ From probabilities to decisions: threshold routing](#routing)
4. [ Composite scoring: weights live in your code](#composite)
5. [ Measuring calibration instead of trusting it](#measuring)
6. [ What the promise does not cover](#edges)
7. [ Knowledge check](#check)
8. [ Further reading](#further)

## What the numbers promise

Jev's real product claim is not speed. It is **calibration**: higher confidence should track higher accuracy, and the model is trained to make that true.

> 🏷️ New term: **Calibration** [glossary](/courses/jev-fundamentals/glossary/)
> A judge is calibrated when its confidence numbers track actual correctness: among the answers it scores at 0.8, roughly 80 percent are actually right. Calibration is what makes confidence usable as a probability, a number you can build thresholds on.

Why would a model be calibrated at all? Because it was trained to be. Jev's training objective is **RLCD** (Reinforcement Learning for Calibrated Decisions), which rewards honest probabilities instead of confident-sounding ones (vendor's framing, launch post). It targets a known mainstream defect: RLHF demonstrably hurts calibration (the GPT-4 technical report's reliability plots show it), and a judge that is confidently wrong is worse than one that hedges honestly. The idea that a model can be trained to know what it does not know predates Jev: Kadavath et al. established it in 2022 (further reading below).

> 🏷️ New term: **RLCD** [glossary](/courses/jev-fundamentals/glossary/)
> Reinforcement Learning for Calibrated Decisions, the training objective that rewards honest probabilities instead of confident-sounding ones.

Here is why this matters more than accuracy. Consider a judge that automates a decision: "act when confidence ≥ 0.9". If the model is calibrated, that line means "act when it is right at least 90 percent of the time", and you can tune it with data. If the model is _miscalibrated_, its 0.9 might be right only 60 percent of the time, and your threshold enforces a policy that does not exist. High accuracy with broken calibration still breaks threshold-based automation, because the number your code acts on stops meaning what you think it means. That is why a miscalibrated judge is worse than its accuracy suggests.

> 💡 Key insight Calibration is a measurable property. Measure it instead of trusting it: a model that is accurate but miscalibrated fails the judge use case even though its verdicts look right. Section 5 shows the two standard instruments.

## Reading a Noul against thresholds

A `noul` answer is a probability with no separate confidence field: the value itself is the uncertainty signal. Reading it against a threshold policy turns it into a verdict:

| Noul value | What it means | Reasonable action |
| --- | --- | --- |
| ≥ 0.95 | Confidently true | Count as `met`, move on |
| 0.70 – 0.95 | Probably true | `met` under a precision-biased threshold of 0.7 |
| 0.30 – 0.70 | Genuinely uncertain | Treat as the abstention band; route to review |
| 0.05 – 0.30 | Probably false | `unmet` |
| ≤ 0.05 | Confidently false | `unmet`, no review needed |

> 🏷️ New term: **Abstention band** [glossary](/courses/jev-fundamentals/glossary/)
> The Noul range 0.30–0.70, where the model is torn between the two answers. The fraction of answers landing there is worth reporting: it is the share of work a human still has to look at.

The split the table enforces is the important design move: **the model supplies probabilities, and the threshold policy is yours**. The same assessment can feed a strict policy (act only at ≥ 0.7, encode a "when in doubt, fail" bias) or a lenient one (act at ≥ 0.5, accept more noise). Where the cutoff sits is a product decision you own in code, not prompt wording.

## From probabilities to decisions: threshold routing

One battery of Noul questions plus one severity `score` is enough to route real traffic, and the vendor's guardrails cookbook works the full example (all figures below are from the vendor cookbook, verified 2026-09-23; numbers produced by `jev-1.12` on 2026-08-15). The shape: screen a message with a battery of hazard nouls (jailbreak, harmful request, medical advice, self-harm) and one severity score. Two thresholds do the routing:

- at or above the **action threshold**, the hazard triggers its configured action;
- at or above the lower **review threshold**, the message goes to a human;
- below both, it passes unless another hazard fires.

The severity score has a threshold of its own and can escalate a review into a block. A "policy" is just those numbers under a name, which turns the trade-off into something the application picks rather than inherits:

```
# Same TypeSafe result: jailbreak=0.74, severity=0.51

strict       review >= 0.35  action >= 0.70  ->  block
permissive   review >= 0.35  action >= 0.85  ->  review
```

<svg viewBox="0 0 700 240" style="width:100%;max-width:680px;height:auto;margin:1.5rem auto;display:block;" aria-label="One cached assessment (jailbreak 0.74, severity 0.51) routed by two policies: strict blocks at action threshold 0.70, permissive sends to review at 0.85" xmlns="http://www.w3.org/2000/svg"> <style> .tr-bg { fill: var(--surface, #fff); } .tr-text { fill: var(--ink, #111); font-family: var(--font-ui, sans-serif); font-size: 13px; } .tr-mono { fill: var(--ink, #111); font-family: var(--font-mono, monospace); font-size: 12px; } .tr-small { fill: var(--ink-muted, #666); font-family: var(--font-ui, sans-serif); font-size: 12px; } .tr-box { fill: var(--surface, #fff); stroke: var(--border, #999); stroke-width: 1.5; } .tr-accentbox { fill: var(--accent, #0066cc); fill-opacity: 0.10; stroke: var(--accent, #0066cc); stroke-width: 1.5; } .tr-arrow { stroke: var(--ink-muted, #666); stroke-width: 1.5; fill: none; marker-end: url(#tr-arrowhead); } .tr-arrowhead { fill: var(--ink-muted, #666); } </style> <defs> <marker id="tr-arrowhead" viewBox="0 0 10 6" refX="10" refY="3" markerWidth="10" markerHeight="6" orient="auto"> <polygon points="0 0, 10 3, 0 6" fill="#666" class="tr-arrowhead"></polygon> </marker> </defs> <rect x="0" y="0" width="700" height="240" rx="8" class="tr-bg" fill="#fff"></rect> <rect x="25" y="55" width="175" height="122" rx="6" class="tr-box" fill="#fff" stroke="#999"></rect> <text x="112" y="80" text-anchor="middle" class="tr-text" fill="#111" font-weight="700">ONE ASSESSMENT</text> <text x="40" y="110" class="tr-mono" fill="#111">jailbreak = 0.74</text> <text x="40" y="132" class="tr-mono" fill="#111">severity = 0.51</text> <text x="40" y="154" class="tr-small" fill="#666">cached; the model is</text> <text x="40" y="170" class="tr-small" fill="#666">not called again</text> <rect x="265" y="30" width="170" height="70" rx="6" class="tr-box" fill="#fff" stroke="#999"></rect> <text x="350" y="50" text-anchor="middle" class="tr-text" fill="#111" font-weight="700">strict policy</text> <text x="350" y="68" text-anchor="middle" class="tr-mono" fill="#111">review ≥ 0.35</text> <text x="350" y="84" text-anchor="middle" class="tr-mono" fill="#111">action ≥ 0.70</text> <rect x="265" y="130" width="170" height="70" rx="6" class="tr-box" fill="#fff" stroke="#999"></rect> <text x="350" y="150" text-anchor="middle" class="tr-text" fill="#111" font-weight="700">permissive policy</text> <text x="350" y="168" text-anchor="middle" class="tr-mono" fill="#111">review ≥ 0.35</text> <text x="350" y="184" text-anchor="middle" class="tr-mono" fill="#111">action ≥ 0.85</text> <rect x="505" y="30" width="165" height="70" rx="6" class="tr-accentbox" fill="#0066cc" fill-opacity="0.10" stroke="#0066cc"></rect> <text x="587" y="60" text-anchor="middle" class="tr-text" fill="#111" font-weight="700">block</text> <text x="587" y="82" text-anchor="middle" class="tr-small" fill="#666">refuse the turn</text> <rect x="505" y="130" width="165" height="70" rx="6" class="tr-accentbox" fill="#0066cc" fill-opacity="0.10" stroke="#0066cc"></rect> <text x="587" y="160" text-anchor="middle" class="tr-text" fill="#111" font-weight="700">review</text> <text x="587" y="182" text-anchor="middle" class="tr-small" fill="#666">a human looks</text> <line x1="200" y1="100" x2="261" y2="65" class="tr-arrow" stroke="#666" fill="none"></line> <line x1="200" y1="132" x2="261" y2="165" class="tr-arrow" stroke="#666" fill="none"></line> <line x1="435" y1="65" x2="503" y2="65" class="tr-arrow" stroke="#666" fill="none"></line> <line x1="435" y1="165" x2="503" y2="165" class="tr-arrow" stroke="#666" fill="none"></line> <text x="350" y="228" text-anchor="middle" class="tr-small" fill="#666">the probabilities do not move; the policy decides how much evidence to demand</text> </svg>

The cookbook's own numbers show why the routing layers earn their keep. Under the strict policy, one message asking about a melatonin dose landed at `medical_advice=0.55` and went to a human (a dosage question mild enough to hand to a person), while an 800 mg ibuprofen dosage reply crossed the severity block line at 2.02 and became a block. A message dressed up as a medical accommodation scored `jailbreak=0.74`: the strict policy blocked it, the permissive one would only review it. Same probabilities, different decisions. The assessment is reusable; only the thresholds change.

> 💡 Key insight TypeSafe supplies the assessment; your application owns the decision. The four-way route (pass, review, block, support) is what a plain block cannot do: it is the difference between helping someone and hanging up on them. Both the hazard battery and the thresholds are yours to edit.

### Uncertainty gates: when the answer is not settled

The thresholds above assume a Noul, whose single probability is both the evidence and the uncertainty statement — the abstention band (0.30–0.70) falls out of it for free. Choice and Score carry two dials instead (S01E01), so "in the band" becomes a predicate — a small OR of conditions, any one of which routes to review:

```
// Choice: margin = winner − runner-up; both gates needed
inBand = confidence < minConfidence  ||  margin < minMargin

// Score: relative to the decision, not intrinsic
inBand = confidence < minConfidence      // hedged position
      || tornDistribution               // mass split across adjacent levels
      || nearThreshold(threshold)       // straddling the tripwire
```

> 🏷️ New term: **margin** [glossary](/courses/jev-fundamentals/glossary/)
> The winner's probability minus the runner-up's. The coin-flip detector for a Choice: the answer always looks like an answer, but a 0.51/0.49 winner is arbitrary between ties. Margin is what tells a decisive 0.85/0.15 from a torn 0.51/0.49 with the same winner label.

The two Choice gates catch different shapes of unsettled: `0.9/0.05/0.05` has a clear winner but low confidence; `0.51/0.49` may carry decent confidence on the winner while the top two sit in a dead heat. Margin catches the second, confidence the first — you need both. For a Score, "in the band" is relative to the decision, not intrinsic: 1.4 is unsettled only if your threshold is nearby, and a Score at 1.9 against a 2.0 block line with torn mass is not a pass — it is a maybe, and maybes go to review.

> 💡 Probabilities choose the action; confidence chooses whether to automate it Because confidence is calibrated, low confidence is a prediction of error, not noise: when the model says 0.55, it is wrong about 45% of the time. So the working rule is two dials in two roles — the probabilities (or the Noul value) decide what to do; the confidence (and the margin) decide whether that action is safe to automate. An unsettled answer has three exits, in order of preference: route to review, take the safe default (the cheaper-to-fail option), or rewrite the question. A recurring tie or a wide band is not the model's indecision — it is the menu, the levels, or the criteria wording, and the tie rate is the metric that measures your schema. Never "interpret harder": act, review, or rewrite.

## Composite scoring: weights live in your code

Routing handles one binary outcome at a time. Ranking several items needs more, and the composite scoring pattern (vendor pattern doc, verified 2026-09-23) is the designed shape for it: break the judgment into independent dimensions, score each separately, and combine with weights you control in code.

The vendor's worked example ranks job candidates on four `score` questions in one request: Python depth, team leadership, system design, generalist range. Each dimension has written levels from "no experience mentioned" to "deep expertise". Then, entirely in code:

```
py      = response.answers["python_depth"].score / 4
lead    = response.answers["team_leadership"].score / 4
arch    = response.answers["system_design"].score / 4
general = response.answers["generalist"].score / 4

# Senior IC
ic_score = (0.40 * py) + (0.10 * lead) + (0.40 * arch) + (0.10 * general)

# Engineering Manager
em_score = (0.15 * py) + (0.40 * lead) + (0.20 * arch) + (0.25 * general)
```

Three properties make the pattern work:

- **Normalization is yours.** Each score is divided by 4 (the highest level index) to land in 0–1 before weighting.
- **Weights are configuration, not prompts.** The same four answers serve two different roles (senior IC, engineering manager) under two weight sets. Changing the ranking logic never touches the model.
- **The arithmetic is visible.** You can see exactly how a final score was produced, and adjust weights when the top-ranked candidates do not match expectations.

This is also where the independence property from Lesson 1 pays off: the four dimensions are scored in parallel in one request, none influenced by another's answer, so a resume that games the leadership dimension does not move the Python score.

> 🔧 What composite scoring does not give you Composition multiplies coverage; it does not create independence between models. Ten questions share one model's blind spots, so the weights combine correlated errors as readily as signals. If the ranking matters, hold back a labeled sample and check the composite against it before you trust the ordering (Lesson 1's "correlated mistakes" edge, restated where it bites).

## Measuring calibration instead of trusting it

Two standard instruments turn "should be calibrated" into a number you can act on:

> 🏷️ New term: **Reliability curve** [glossary](/courses/jev-fundamentals/glossary/)
> Predicted probability buckets against observed accuracy. Bucket answers by stated confidence, measure how often each bucket was right, and plot the two against each other. The diagonal is perfect calibration.

> 🏷️ New term: **Brier score** [glossary](/courses/jev-fundamentals/glossary/)
> A proper scoring rule: the mean squared difference between stated probability and outcome. It punishes confident wrong answers hardest, so a well-calibrated model scores lower than a lucky one.

The reliability curve shows you _where_ the model is off (a bucket that says 0.9 but is right 70 percent of the time); the Brier score compresses the whole curve into one comparable number. Measure both against human labels on your own inputs, not the vendor's benchmarks: calibration claims are trained behavior, and your input mix is the distribution that matters.

One more instrument belongs in the set when the task is agreement with human judgment: **κ (Cohen's kappa)**, agreement between two raters beyond what chance would produce. It is the standard metric for "does the model's verdict match a careful human's verdict" on binary met/unmet decisions, and it is what a judge's calibration monitoring should track over time.

> 🏷️ New term: **κ (Cohen's kappa)** [glossary](/courses/jev-fundamentals/glossary/)
> Rater agreement beyond chance. Chance-adjusted, so it does not reward agreement that mere class balance would produce.

> 🔧 The operating loop Set the threshold policy from labeled examples → run the judge → track κ and the reliability curve per run → when your input mix changes, expect the numbers to move and re-check before acting on them. Calibration is trained behavior plus your distribution; both halves need watching.

## What the promise does not cover

Three edges keep the calibration claim honest:

### The abstention band is a design decision

Everything between 0.30 and 0.70 is genuinely uncertain, and automation has to decide what happens there. The guardrails pattern routes that band to human review; a forced-choice schema with no `other` option instead forces an answer even when the true answer is missing from the menu. The schema author carries that burden (Lesson 1), and the abstention-band fraction is the metric that tells you how much of your traffic is quietly waiting for a human.

### Calibration claims are vendor-measured

The RLCD story, the latency claims, and the headline eval numbers all come from the vendor's own workflows. Treat them as ceilings, and measure on your own labels. A vendor that trains for calibration has an incentive to be right about it, but your distribution is not their benchmark.

> 🏷️ New term: **OOD drift** [glossary](/courses/jev-fundamentals/glossary/)
> Out-of-Distribution drift: inputs and policies shift over time, and probabilities that were reliable last month can move without notice. The countermeasure is the monitoring loop from section 5: κ and reliability curves, tracked on every run.

### Confidence is per answer, not per request

Each question in a request carries its own probabilities. A batch of five answers is five separate calibration events, not one confident verdict; report the band or the minimum confidence across the battery, not an average that hides the one answer the model was torn on.

## Knowledge check

**What does calibration mean for a judge model?**

- **✓** Confidence tracks actual accuracy
- Confidence stays close to 1.0
- Probabilities sum to exactly 2
- Scores match human rankings

Calibration is the link between stated confidence and real correctness: when the model says 0.9, it should be right about 90 percent of the time. RLCD trains for this directly.

**Why does a miscalibrated judge break threshold-based automation, even with high accuracy?**

- **✓** Thresholds act on the number, so the number must mean what it claims
- High accuracy requires probabilities near 1.0
- Miscalibrated models refuse to answer
- Thresholds only work on choice questions

Your code acts on the stated number. If 0.9 is really right 60 percent of the time, the action threshold enforces a policy that does not exist. Accuracy alone does not repair that; the number itself is wrong.

**In the guardrails routing pattern, what decides between pass, review, block, and support?**

- A separate model call per outcome
- **✓** Thresholds in your code over the same probabilities
- The confidence field alone
- A random draw weighted by severity

One cached assessment feeds the threshold policy in code: hazards above the action threshold trigger their action, above the review threshold they go to a human. Changing the policy re-routes without calling the model again.

**A Choice comes back 0.51/0.49 with decent confidence on the winner. What does good decision code do?**

- Act on the winner — it is the top answer
- **✓** Route to review: the thin margin makes the winner a coin flip between ties
- Average the two probabilities first
- Re-ask the same question until the tie breaks

The margin (winner minus runner-up) is near zero, so the winner is arbitrary between ties — acting on it silently converts "I don't know" into a decision. Margin and confidence catch different shapes of unsettled, and any one firing routes to review. A recurring tie is a schema problem: merge the overlapping options or add an escape-hatch option.

**What does the composite scoring pattern do?**

- Trains one model per criterion
- Merges all items into one call
- **✓** Weights independent scores in code
- Drops questions below 0.5

Break the judgment into atomic questions, then combine with weights in your code. The weights live in code, so ranking logic changes without touching a prompt.

## Further reading

- [_Composite scoring pattern (TypeSafe AI docs)_](https://docs.typesafe.ai/patterns/composite-scoring): the full resume-screening worked example this lesson's section 4 compresses, with runnable weights.
- [_Guardrails for LLMs cookbook (TypeSafe AI)_](https://docs.typesafe.ai/cookbooks/llm_guardrails): the complete threshold-routing worked example: hazard nouls, severity score, two named policies, and every sample message's routing result.
- [_Language Models (Mostly) Know What They Know_ (Kadavath et al., 2022)](https://arxiv.org/abs/2207.05221): the research result behind "calibration is trainable", the idea RLCD productizes.
- [_Introducing System One Models and Jev_](https://typesafe.ai/blog/introducing-system-one-models-and-jev): the launch post, for the RLCD framing in the vendor's own words.
- [_GPT-4 Technical Report_ (OpenAI, 2023)](https://arxiv.org/abs/2303.08774): the reliability plots showing RLHF's calibration cost, the mainstream defect RLCD targets.
- [_Jev: The Language Model That Won't Talk_ (Anthony Maio)](https://anthonymaio.substack.com/p/jev-the-language-model-that-wont): the calibration caveats and the case for treating confidence numbers with care.

> **🧭 Practice** Your turn. Take one binary decision you already automate (or would like to) and write its threshold policy down: what noul value counts as met, what band routes to review, and what the action threshold encodes about your error costs. Then open the TypeSafe Playground, run one real input through it, and check whether your thresholds would have routed the input the way you intended. If an AI assistant is part of your workflow, Copy page in the header menu hands it this whole lesson as markdown, citation included. Ask it to quiz you on the knowledge check, or to explain a section a different way. The glossary is the canonical vocabulary when terms blur.

## See also

- [📖 Jev Glossary](/courses/jev-fundamentals/glossary/)
- [📐 S01E01: Jev answers instead of prose](/courses/jev-fundamentals/s01e01-jev-answers-instead-of-prose/)
---

Source: [Confidence, Calibration & Composite Scoring](https://peripatos.dev/courses/jev-fundamentals/s01e03-jev-confidence-calibration-composite/), Peripatos — free to read; reuse by permission.
