S01E01: Jev — The Model That Answers Instead of Writes
Jev does one job: it looks at a situation you hand it and returns typed, calibrated decisions your code can use directly: no prose to parse, no tokens to recover values from. By the end of this lesson you can read a Jev request and a Jev answer as fluently as you read a function signature.
Why Jev exists: software wants decisions
When you ask a language model a question, it answers in prose. Even when you beg it for JSON, the answer is still generated token by token, and your code has to recover a value from that string: parse it, validate it, handle the case where the string is almost-JSON. The model writes; you dig the decision out of the writing.
Jev (TypeSafe AI, launched September 2026) deletes that last step. It is a System One model, named after Kahneman's fast, intuitive judgment mode. It gives up text generation entirely. It reads a state (the context you supply), answers a fixed set of typed questions about it, and returns values your code can use directly.
| LLM (System Two) | Jev (System One) | |
|---|---|---|
| Output | Generated text | Typed values + probabilities |
| Failure mode | Wrong or unparseable prose | Wrong but well-typed answer |
| Latency source | Token count (seconds to minutes for frontier models) | Single forward pass (70–500 ms claimed) |
| Cost source | Input + output tokens | Input tokens only (output free) |
| Best at | Writing, reasoning, synthesis | Classify, route, score, verify |
The unit of work is a judgment a knowledgeable person makes in a second given the right context. Not a paragraph. Not a conversation. One glance, one typed answer. If a task reduces to many such judgments, Jev is in scope.
The three primitives: Noul, Choice, Score
Every Jev question is one of three types. Together they cover nearly everything code asks about content:
| Type | Reach for it when |
|---|---|
noul | Binary checks: "does this message satisfy the criterion?" |
choice | Routing: "which topic?", "which team handles this?" |
score | Ranked judgment: "how frustrated?", "how severe?" |
A real request
POST https://api.typesafe.ai/v1/systemone
Authorization: Bearer $TYPESAFE_API_KEY
{
"state": "Hi, I've been trying to connect my Stripe account
for 3 days and the integration keeps failing.
I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated the customer appears",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language"
]
},
"is_urgent": {
"type": "noul",
"instructions": "The message conveys urgency or time-sensitivity"
}
}
}
Three ideas to notice in the request body:
state: one shared context. Every question in the request sees this exact text and nothing else.questions: a keyed dictionary. You name each question; answers come back under the same names.criteria: you write the option descriptions or score levels. The model never invents labels.
The three type values in that request (choice, score, noul) are the primitives. One at a time:
The probability that a statement is true, from 0 to 1. Near 0.5 means the model is genuinely uncertain.
One picked option from a menu you define at request time, plus the full probability distribution over every option.
A position on ordered levels you define, 0-indexed, and it can land between levels when the model is torn.
You define the answer space at request time; the model has no fixed labels: the same "runtime-defined output space" pattern as zero-shot classification, but with LLM-grade world knowledge behind it.
Anatomy of a request and an answer
Reading an answer:
choice: the winner.probabilitiesis the full distribution over every option (it sums to 1).confidenceis the model's certainty in the winner.score: the position on your levels, 0-indexed, plus the same probability machinery.noul: just the probability. There is no separate confidence: the value itself is the uncertainty signal. 0.5 means the model honestly does not know.
The response carries two uncertainty numbers, and they answer different questions. The probabilities compare the options against each other: technical's 0.85 is its share of the menu. The confidence says how much to trust the pick itself: 0.85 here. Same value in this example, but they can diverge — 0.85/0.12/0.03 with confidence 0.55 means the winner is clear yet the model does not stand behind it.
Torn answers are the case worth training on. A Choice whose top two options split 0.51/0.49 is a coin flip wearing the winner's label: the winner is still returned, but it is arbitrary between ties, and nothing in the value itself announces this — your code has to compute it (winner minus runner-up; Lesson 3's margin). A torn Score works the same way: mass split across adjacent levels makes the position land between levels, which is information rather than loss. And a Noul needs none of this machinery — its single probability already carries the uncertainty, which is exactly why it has no confidence field.
Questions inside one request are independent by design: no answer ever becomes context for another. This is not a limitation to work around; it is the property that makes the parallel fan-out cheap and the per-question blast radius small. When judgments do depend on each other, you make a second request, which the vendor documents as "the exception, not the rule".
One request, many questions
Because questions are independent, batching is the designed shape, not an optimization. The vendor cookbook measured batching 13 questions into one call at 12.2× cheaper and 10.0× faster than 13 separate calls, with identical answers (vendor cookbook, verified 2026-09-23). The vendor claims a large shared state context and says capacity limits "adjust dynamically"; treat both as ceilings and measure, don't design against them.
The design philosophy goes further: ask speculatively. Questions are cheap and parallel, so you ask the ones you might need and ignore the answers you do not use. The composite scoring pattern builds on this: break a complex judgment into independent atomic scores, then combine them with weights you control in code. The weights live in your code, not in the model's head: ranking logic changes without touching a prompt. Lesson 3 works a full composite-scoring example end to end.
One request per input bounds the damage one malicious input can do to its own verdict. Within that request, the questions are supposed to be independent: a hostile message that games the urgency question does not thereby game the topic question. Independence by design is what makes per-item isolation and batching compose cleanly.
From a rubric to questions: a worked example
The pattern above becomes concrete when you already have a rubric: a fixed list of criteria a reviewer checks, one binary verdict each. Here is a self-contained one. Suppose you screen incoming feature requests for a small product team, and every request is scored against four criteria:
| Criterion | What it asks | Behavioral anchors |
|---|---|---|
real_problem | Does the request describe a problem the submitter actually has? | true: names a concrete, recurring workflow failure. false: a wish list item ("it would be cool if…") |
specific | Is the request specific enough to act on? | true: names the feature, the context, and the expected outcome. false: vague ("make it better") |
in_scope | Does it fit the product's actual audience? | true: matches the product's documented use cases. false: asks for a different product |
new_information | Does it add something past requests missed? | true: gives a new use case, data point, or workaround. false: restates an existing request verbatim |
Each criterion becomes one noul question, and the behavioral anchors become its true/false criteria. One request covers all four:
{
"state": "Feature request: 'Every time I export a report on
mobile the column widths reset. I export 20+ reports
a week for client meetings and have to fix the
widths by hand each time. Please persist column
widths per report template.'",
"model": "jev-latest",
"questions": {
"real_problem": {
"type": "noul",
"instructions": "Does this request describe a problem the submitter actually experiences?",
"criteria": {
"true": "Names a concrete, recurring problem in their own workflow.",
"false": "A wish list item with no concrete problem behind it."
}
},
"specific": {
"type": "noul",
"instructions": "Is the request specific enough to act on?",
"criteria": {
"true": "Names the feature, the context, and the expected behavior.",
"false": "Vague or general, with no concrete behavior requested."
}
},
"in_scope": {
"type": "noul",
"instructions": "Does the request fit the product's actual audience and use cases?",
"criteria": {
"true": "Matches the product's documented use cases.",
"false": "Asks for something a different product would provide."
}
},
"new_information": {
"type": "noul",
"instructions": "Does the request add something earlier requests missed?",
"criteria": {
"true": "Contributes a new use case, data point, or workaround.",
"false": "Restates what existing requests already cover."
}
}
}
}
The mapping is mechanical, and that is the point:
| Rubric element | Jev element | Notes |
|---|---|---|
Criterion (e.g. specific) | One noul question: "Does this item satisfy: <criterion>?" | Four questions, one request |
| Behavioral anchors | criteria: {true: "…", false: "…"} | Anchors become the option definitions |
| met / unmet verdict | Threshold on the noul value | ≥ 0.5 default; raise it to encode a "when in doubt, fail" bias (tune on labeled examples) |
| Final score | Computed in code, from verdicts | Score = met count / 4; pass at a threshold you own |
| Written justification per verdict | Dropped (Jev does not write prose) | The trace loses the "why"; route on confidence bands instead |
| One call per item | One request per item | Blast radius stays bounded per item |
One honest loss in this translation: the written justification. A classic judge trace explains why an item failed, which is half of what makes verdicts reviewable. With Jev you get a probability instead of an argument: the review queue reads confidence bands, not prose. That trade is measurable: compare how much review time the explanations actually save against how well the probabilities predict review outcomes.
Accuracy is not the only thing to measure when you swap a judge. Jev is a different architecture trained for a different output space. If its errors land in different places than your current LLM judge's errors, a panel of both is more independent than two same-family models, which is the entire argument for judge-panel diversity. Measure the error overlap before you trust the panel.
What Jev is not
The claims need their honest edges before you build on them:
Type safety is a shape guarantee, not a truth guarantee
"Can't hallucinate" means the answer's shape always matches the schema: the value is well-typed by construction. The model can still pick a wrong-but-well-typed option. A forced-choice schema also cannot abstain when the true answer is missing from the menu. Canonical counterexample: "does the user want a human agent?" gets a forced met/unmet even when the true answer is "yes, but not right now". The vendor's guidance (add an other or none of the above option) puts the abstention burden on the schema author: you.
Correlated mistakes survive composition
Decomposing one complex judgment into ten Jev questions shares one model's blind spots across all ten. Composition multiplies coverage; it does not create independence. This is the same lesson as judge panels: compare a judge against human labels, never against another model alone.
Practical edges
- Closed model: no open weights, no local deployment. Private context leaves your machine.
- Vendor pricing ($0.042/1M input tokens as of 2026-09-23) may be subsidized; sustainability is unproven.
- Headline eval numbers are vendor-run on vendor-authored workflows; treat them as ceilings.
Out-of-Distribution drift: inputs and policies shift over time, and probabilities that were reliable last month can move without notice. The countermeasure is calibration monitoring: track agreement with human labels on every run, and re-check the numbers when your input mix changes.
Knowledge check
📖 This knowledge check is an interactive quiz (6 questions). Open this page in a browser to take it; reader mode shows only the prose.
Further reading
- Primitives (TypeSafe AI docs): the concrete API surface this lesson compresses: question types, response fields, fan-out economics, abstention guidance. Read it once, then keep it open while writing your first request.
- Introducing System One Models and Jev: the vendor's own framing, claims, and the Jevons-paradox thesis. Read it alongside the skeptic case.
- Jev: The Language Model That Won't Talk (Anthony Maio): the best critical analysis: the hidden tax of strings, calibration caveats, and the schema-author burden. Steelman before you evaluate.
Your turn. Open the TypeSafe Playground, paste one real message from your own work as the state, and add four Noul questions: one per criterion from a checklist you already use, with the behavioral anchors as the true/false criteria. Compare the four probabilities with what your own gut said about the same message.
If an AI assistant is part of your workflow, Copy page in the header menu hands it this whole lesson as markdown, citation included. Ask it to quiz you on the knowledge check, or to explain a section a different way. The glossary is the canonical vocabulary when terms blur.