S01E01: Jev — The Model That Answers Instead of Writes

  • Level L1 (beginner)
  • no prerequisites|~15 min|verified 2026-09-24
  • system-one-modelsjevstructured-outputsclassificationllm-as-judge

Jev does one job: it looks at a situation you hand it and returns typed, calibrated decisions your code can use directly: no prose to parse, no tokens to recover values from. By the end of this lesson you can read a Jev request and a Jev answer as fluently as you read a function signature.

Why Jev exists: software wants decisions

When you ask a language model a question, it answers in prose. Even when you beg it for JSON, the answer is still generated token by token, and your code has to recover a value from that string: parse it, validate it, handle the case where the string is almost-JSON. The model writes; you dig the decision out of the writing.

Jev (TypeSafe AI, launched September 2026) deletes that last step. It is a System One model, named after Kahneman's fast, intuitive judgment mode. It gives up text generation entirely. It reads a state (the context you supply), answers a fixed set of typed questions about it, and returns values your code can use directly.

LLM (System Two)Jev (System One)
OutputGenerated textTyped values + probabilities
Failure modeWrong or unparseable proseWrong but well-typed answer
Latency sourceToken count (seconds to minutes for frontier models)Single forward pass (70–500 ms claimed)
Cost sourceInput + output tokensInput tokens only (output free)
Best atWriting, reasoning, synthesisClassify, route, score, verify
💡 Key insight

The unit of work is a judgment a knowledgeable person makes in a second given the right context. Not a paragraph. Not a conversation. One glance, one typed answer. If a task reduces to many such judgments, Jev is in scope.

The three primitives: Noul, Choice, Score

Every Jev question is one of three types. Together they cover nearly everything code asks about content:

TypeReach for it when
noulBinary checks: "does this message satisfy the criterion?"
choiceRouting: "which topic?", "which team handles this?"
scoreRanked judgment: "how frustrated?", "how severe?"

A real request

POST https://api.typesafe.ai/v1/systemone
Authorization: Bearer $TYPESAFE_API_KEY

{
  "state": "Hi, I've been trying to connect my Stripe account
            for 3 days and the integration keeps failing.
            I'm losing sales. Please help ASAP.",
  "model": "jev-latest",
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this",
      "criteria": {
        "billing":   "Payment or subscription issues",
        "technical": "Bugs or integration problems",
        "sales":     "Pricing or account questions"
      }
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated the customer appears",
      "criteria": [
        "Calm, just stating facts",
        "Frustrated but civil",
        "Very angry, strong language"
      ]
    },
    "is_urgent": {
      "type": "noul",
      "instructions": "The message conveys urgency or time-sensitivity"
    }
  }
}

Three ideas to notice in the request body:

  • state: one shared context. Every question in the request sees this exact text and nothing else.
  • questions: a keyed dictionary. You name each question; answers come back under the same names.
  • criteria: you write the option descriptions or score levels. The model never invents labels.

The three type values in that request (choice, score, noul) are the primitives. One at a time:

🏷️ New termNoul· glossary

The probability that a statement is true, from 0 to 1. Near 0.5 means the model is genuinely uncertain.

🏷️ New termChoice· glossary

One picked option from a menu you define at request time, plus the full probability distribution over every option.

🏷️ New termScore· glossary

A position on ordered levels you define, 0-indexed, and it can land between levels when the model is torn.

You define the answer space at request time; the model has no fixed labels: the same "runtime-defined output space" pattern as zero-shot classification, but with LLM-grade world knowledge behind it.

Anatomy of a request and an answer

REQUEST state: "…" questions: { department: choice frustration: score is_urgent: noul Jev RESPONSE answers: { department: "technical" conf 0.85 · probs 0.85/0.12/0.03 frustration: 1.0 conf 1.0 · probs 0/1/0 is_urgent: 1.0 (noul) one request · questions evaluated independently and in parallel usage.input_tokens counts; output is free

Reading an answer:

  • choice: the winner. probabilities is the full distribution over every option (it sums to 1). confidence is the model's certainty in the winner.
  • score: the position on your levels, 0-indexed, plus the same probability machinery.
  • noul: just the probability. There is no separate confidence: the value itself is the uncertainty signal. 0.5 means the model honestly does not know.
💡 The two dials: probabilities and confidence

The response carries two uncertainty numbers, and they answer different questions. The probabilities compare the options against each other: technical's 0.85 is its share of the menu. The confidence says how much to trust the pick itself: 0.85 here. Same value in this example, but they can diverge — 0.85/0.12/0.03 with confidence 0.55 means the winner is clear yet the model does not stand behind it.

Torn answers are the case worth training on. A Choice whose top two options split 0.51/0.49 is a coin flip wearing the winner's label: the winner is still returned, but it is arbitrary between ties, and nothing in the value itself announces this — your code has to compute it (winner minus runner-up; Lesson 3's margin). A torn Score works the same way: mass split across adjacent levels makes the position land between levels, which is information rather than loss. And a Noul needs none of this machinery — its single probability already carries the uncertainty, which is exactly why it has no confidence field.

💡 Key insight

Questions inside one request are independent by design: no answer ever becomes context for another. This is not a limitation to work around; it is the property that makes the parallel fan-out cheap and the per-question blast radius small. When judgments do depend on each other, you make a second request, which the vendor documents as "the exception, not the rule".

One request, many questions

Because questions are independent, batching is the designed shape, not an optimization. The vendor cookbook measured batching 13 questions into one call at 12.2× cheaper and 10.0× faster than 13 separate calls, with identical answers (vendor cookbook, verified 2026-09-23). The vendor claims a large shared state context and says capacity limits "adjust dynamically"; treat both as ceilings and measure, don't design against them.

The design philosophy goes further: ask speculatively. Questions are cheap and parallel, so you ask the ones you might need and ignore the answers you do not use. The composite scoring pattern builds on this: break a complex judgment into independent atomic scores, then combine them with weights you control in code. The weights live in your code, not in the model's head: ranking logic changes without touching a prompt. Lesson 3 works a full composite-scoring example end to end.

🔧 Why the blast radius stays small

One request per input bounds the damage one malicious input can do to its own verdict. Within that request, the questions are supposed to be independent: a hostile message that games the urgency question does not thereby game the topic question. Independence by design is what makes per-item isolation and batching compose cleanly.

From a rubric to questions: a worked example

The pattern above becomes concrete when you already have a rubric: a fixed list of criteria a reviewer checks, one binary verdict each. Here is a self-contained one. Suppose you screen incoming feature requests for a small product team, and every request is scored against four criteria:

CriterionWhat it asksBehavioral anchors
real_problemDoes the request describe a problem the submitter actually has?true: names a concrete, recurring workflow failure. false: a wish list item ("it would be cool if…")
specificIs the request specific enough to act on?true: names the feature, the context, and the expected outcome. false: vague ("make it better")
in_scopeDoes it fit the product's actual audience?true: matches the product's documented use cases. false: asks for a different product
new_informationDoes it add something past requests missed?true: gives a new use case, data point, or workaround. false: restates an existing request verbatim

Each criterion becomes one noul question, and the behavioral anchors become its true/false criteria. One request covers all four:

{
  "state": "Feature request: 'Every time I export a report on
            mobile the column widths reset. I export 20+ reports
            a week for client meetings and have to fix the
            widths by hand each time. Please persist column
            widths per report template.'",
  "model": "jev-latest",
  "questions": {
    "real_problem": {
      "type": "noul",
      "instructions": "Does this request describe a problem the submitter actually experiences?",
      "criteria": {
        "true":  "Names a concrete, recurring problem in their own workflow.",
        "false": "A wish list item with no concrete problem behind it."
      }
    },
    "specific": {
      "type": "noul",
      "instructions": "Is the request specific enough to act on?",
      "criteria": {
        "true":  "Names the feature, the context, and the expected behavior.",
        "false": "Vague or general, with no concrete behavior requested."
      }
    },
    "in_scope": {
      "type": "noul",
      "instructions": "Does the request fit the product's actual audience and use cases?",
      "criteria": {
        "true":  "Matches the product's documented use cases.",
        "false": "Asks for something a different product would provide."
      }
    },
    "new_information": {
      "type": "noul",
      "instructions": "Does the request add something earlier requests missed?",
      "criteria": {
        "true":  "Contributes a new use case, data point, or workaround.",
        "false": "Restates what existing requests already cover."
      }
    }
  }
}

The mapping is mechanical, and that is the point:

Rubric elementJev elementNotes
Criterion (e.g. specific)One noul question: "Does this item satisfy: <criterion>?"Four questions, one request
Behavioral anchorscriteria: {true: "…", false: "…"}Anchors become the option definitions
met / unmet verdictThreshold on the noul value≥ 0.5 default; raise it to encode a "when in doubt, fail" bias (tune on labeled examples)
Final scoreComputed in code, from verdictsScore = met count / 4; pass at a threshold you own
Written justification per verdictDropped (Jev does not write prose)The trace loses the "why"; route on confidence bands instead
One call per itemOne request per itemBlast radius stays bounded per item

One honest loss in this translation: the written justification. A classic judge trace explains why an item failed, which is half of what makes verdicts reviewable. With Jev you get a probability instead of an argument: the review queue reads confidence bands, not prose. That trade is measurable: compare how much review time the explanations actually save against how well the probabilities predict review outcomes.

🎯 The deepest test

Accuracy is not the only thing to measure when you swap a judge. Jev is a different architecture trained for a different output space. If its errors land in different places than your current LLM judge's errors, a panel of both is more independent than two same-family models, which is the entire argument for judge-panel diversity. Measure the error overlap before you trust the panel.

What Jev is not

The claims need their honest edges before you build on them:

Type safety is a shape guarantee, not a truth guarantee

"Can't hallucinate" means the answer's shape always matches the schema: the value is well-typed by construction. The model can still pick a wrong-but-well-typed option. A forced-choice schema also cannot abstain when the true answer is missing from the menu. Canonical counterexample: "does the user want a human agent?" gets a forced met/unmet even when the true answer is "yes, but not right now". The vendor's guidance (add an other or none of the above option) puts the abstention burden on the schema author: you.

Correlated mistakes survive composition

Decomposing one complex judgment into ten Jev questions shares one model's blind spots across all ten. Composition multiplies coverage; it does not create independence. This is the same lesson as judge panels: compare a judge against human labels, never against another model alone.

Practical edges

  • Closed model: no open weights, no local deployment. Private context leaves your machine.
  • Vendor pricing ($0.042/1M input tokens as of 2026-09-23) may be subsidized; sustainability is unproven.
  • Headline eval numbers are vendor-run on vendor-authored workflows; treat them as ceilings.
🏷️ New termOOD drift· glossary

Out-of-Distribution drift: inputs and policies shift over time, and probabilities that were reliable last month can move without notice. The countermeasure is calibration monitoring: track agreement with human labels on every run, and re-check the numbers when your input mix changes.

Knowledge check

📖 This knowledge check is an interactive quiz (6 questions). Open this page in a browser to take it; reader mode shows only the prose.

Further reading

🧭 Practice

Your turn. Open the TypeSafe Playground, paste one real message from your own work as the state, and add four Noul questions: one per criterion from a checklist you already use, with the behavioral anchors as the true/false criteria. Compare the four probabilities with what your own gut said about the same message.

If an AI assistant is part of your workflow, Copy page in the header menu hands it this whole lesson as markdown, citation included. Ask it to quiz you on the knowledge check, or to explain a section a different way. The glossary is the canonical vocabulary when terms blur.

See also