📚 Course home · llms.txt (lesson order, glossary, and citation terms).
Jev Glossary
Reference for the Jev course; figures verified 2026-09-23. The compressed vocabulary below is what every lesson reads from.
The model
- System One model — gives up text generation; returns typed, calibrated decisions about a supplied state (Kahneman’s fast judgment as an API)
- state — the shared context one request carries; every question sees exactly this text
- question / answer — one typed ask (noul, choice, score); answers keyed by question name
- criteria — the option/level/true-false definitions you author at request time; the model never invents labels
- type safety — shape guarantee, not a truth guarantee
The primitives
- Noul — probability a statement is true, 0–1; no separate confidence; near 0.5 = does not know
- Choice — one picked option from your option map, plus the full distribution
- Score — position on your ordered levels, 0-indexed; can land between levels
- confidence — certainty in the top answer; the calibration promise: tracks actual accuracy
- probabilities — full distribution over all options, sums to 1
- forced choice / abstention — a menu forces an answer even when the true answer is absent; your schema creates the escape hatch
Calibration and measurement
- calibration — stated probability matches observed accuracy; measured with a reliability curve
- RLCD — Reinforcement Learning for Calibrated Decisions; the vendor’s training objective for Jev
- abstention band — Noul 0.30–0.70; the uncertainty range routed to human review; the fraction landing there is a reported metric. For Choice and Score the equivalent is a predicate: low confidence, thin margin, or a torn distribution near a threshold — any one routes to review
- margin — a Choice’s winner probability minus the runner-up’s; the coin-flip detector. A decisive 0.85/0.15 and a torn 0.51/0.49 share the same winner label; the margin tells them apart. A recurring tie is a schema metric: the menu overlaps or lacks an escape hatch
- torn answer — an answer whose distribution is split across options (Choice near-tie) or adjacent levels (Score landing between levels): the model is genuinely unsettled, the top answer is arbitrary, and the routing decision belongs to review or the safe default
- reliability curve — predicted probability buckets vs observed accuracy; diagonal = perfectly calibrated
- Brier score — proper scoring rule: mean squared difference between stated probability and outcome; punishes confident wrong answers
- κ (Cohen’s kappa) — rater agreement beyond chance; the standard metric for judge-vs-human agreement on binary verdicts
- OOD drift — probabilities quietly shift as inputs change; countermeasure: periodic re-measurement on labeled samples
Patterns
- parallel fan-out — many independent questions in one request (13→1: 12.2× cheaper, 10.0× faster, vendor cookbook verified 2026-09-23)
- composite scoring — independent atomic questions combined with weights in code
- threshold routing — the same probabilities feed pass/review/block/support decisions under thresholds your code owns (vendor guardrails cookbook)
- capture-first — cache every raw answer before any analysis, so analysis can improve later without re-running the model
- interleaved timing (ABAB) — alternate the two conditions per measured unit so time-varying noise hits both equally and cancels
- decorrelation — two models making different mistakes; the architecture-diversity hypothesis behind a mixed judge panel
Further reading
If an AI assistant is part of your workflow, the site’s llms.txt lists the
markdown twin of this glossary alongside every lesson; ask it to flashcard you
until every term restates cleanly.