---
title: "Jev State and Typed Questions"
description: "Author the two things you write by hand in every Jev request — the state and the questions. The three question types up close, criteria that define your answer space, independence by design, and a complete worked request."
license: "© 2026 Peripatos — free to read; quoting with credit welcome, reuse by permission. Unofficial educational resource; Jev and TypeSafe are trademarks of TypeSafe AI."
dateModified: 2026-09-23
source: "https://peripatos.dev/courses/jev-fundamentals/s01e02-jev-state-and-typed-questions/"
---

*Level L1 (beginner) — prerequisite: S01E01 · ~15 min · verified 2026-09-24 — Tags: jev, state, typed-questions, criteria, api-design*

A Jev request has two parts you author by hand: the **state** and the **questions**. S01E01 showed both in one paragraph and moved on. This lesson slows down on them, because they are where request quality is actually decided. By the end you can write a complete request: a state that carries the right context, questions that return exactly the values your code wants, and criteria that hold the model to your answer space.

## In this lesson

1. [ The state: one shared context](#state)
2. [ Three question types, up close](#types)
3. [ Criteria: your answer space](#criteria)
4. [ Designing for independence](#independence)
5. [ A complete request: one expense line](#triage)
6. [ Knowledge check](#check)
7. [ Further reading](#reading)

## The state: one shared context

The `state` is the context one request carries. Every question in the request sees exactly this text, and nothing else. That single sentence has three consequences worth pausing on.

First, the state is the _only_ channel in. If a question needs a fact, the fact has to be in the state when the request is written. There is no way for one question to consult another question's answer, and there is no follow-up turn where the model asks you for more context. A request is one shot: whatever a knowledgeable person would need to judge, you put in the state up front.

Second, the state is _shared_. Everything you put there is visible to every question. That is what makes the questions independent and cheap to batch, and it is also a design constraint: question-specific guidance does not belong in the state, because every other question would read it too.

Third, the state is plain text you compose. Same state plus same questions means the model judges the same thing. When a verdict surprises you, the state is the first place to look: what did the model actually see?

> 💡 What belongs in a state The raw material, plus the context a knowledgeable person would need: for an expense line, the memo, the amount, the claimed category, and the policy numbers it must fit under. Instructions do not belong there ("be strict" lives on the questions), and neither does context only one question should weigh. A good test: read the state as if you had to make the judgment yourself in one second. If you would ask "where is…?", it is missing.

## Three question types, up close

S01E01 introduced the three primitives in a table. Here is each one as you actually write it, with the part people get wrong called out.

### Noul: one verifiable statement

```
"is_receipt_referenced": {
  "type": "noul",
  "instructions": "The memo references a receipt or invoice",
  "criteria": {
    "true":  "Names or clearly references a receipt, invoice, or attachment.",
    "false": "No receipt is referenced; the claim rests on the memo alone."
  }
}
```

A noul asks one statement that is either true or false, and returns the probability that it is true. The `criteria` are the true/false definitions, and they are yours to write: "references a receipt" needs a definition, or the model supplies its own. Near 0.5, the model genuinely does not know; what your code does with that band is S01E03's subject.

### Choice: your menu, your escape hatch

```
"lane": {
  "type": "choice",
  "instructions": "Who should handle this expense line",
  "criteria": {
    "auto_approve":   "Within policy limits, routine business expense",
    "manager_review": "Over a limit or policy-adjacent; a manager signs off",
    "finance_audit":  "Possible duplicate, split billing, or odd merchant",
    "other":          "None of the above fits cleanly"
  }
}
```

A choice question picks one option from the map you define and returns the winner, the full probability distribution over every option, and a confidence. The winner is only as good as the menu: if the true answer is not one of the options, the model still has to pick one.

Watch the distribution, not just the winner. A near-tie (`0.51/0.49`) means the model is genuinely torn between two options — the winner is arbitrary between ties, and a recurring tie usually means your menu is wrong: the options overlap (merge them) or the truth lives outside the menu (the `other` lane above). What code does with that is S01E03's margin gate.

> 🏷️ New term: **forced choice / abstention** [glossary](/courses/jev-fundamentals/glossary/)
> A menu forces an answer even when the true answer is absent from it. Your schema creates the escape hatch: an other option, or a "none of the above" lane. If you omit it, the model must force reality into your menu, and the probabilities will tell on it only if you look.

### Score: your ordered levels

```
"policy_risk": {
  "type": "score",
  "instructions": "Risk that this line violates expense policy",
  "criteria": [
    "Routine: well within policy, nothing unusual",
    "Minor: technically claimable, but worth a second look",
    "Serious: likely policy violation, needs a ruling",
    "Severe: clear violation; flag for finance"
  ]
}
```

A score question returns a position on your ordered levels, 0-indexed. The level descriptions are the criteria: they are what separates "minor" from "serious". And the position is not forced onto an integer: the answer can land _between_ levels when the model is torn, which is information your code gets to interpret rather than lose.

The between-levels position is the probability-weighted view of the level distribution: torn roughly evenly between levels 0 and 1 lands near 0.5, and the distribution tells you where the torn-ness sits — between two _adjacent_ levels, not scattered across the scale. Because the score is a summary of the distribution, two different distributions can produce nearly the same number; if your logic depends on between-levels values, read the distribution alongside it.

## Criteria: your answer space

Every label the model can return, you authored at request time. There is no fixed vocabulary behind the API and no label the model invents on its own. The three types express this differently:

| Type | Criteria are | The question they answer |
| --- | --- | --- |
| `noul` | True/false definitions | What exactly counts as true here? |
| `choice` | Option descriptions | What makes this option this option? |
| `score` | Ordered level descriptions | What separates one level from the next? |

> 🏷️ New term: **criteria** [glossary](/courses/jev-fundamentals/glossary/)
> The option, level, or true/false definitions you author at request time. Criteria define the answer space; the model locates the state inside it. Reworded criteria change the question being graded, so treat them like code: versioned, reviewed, and not edited casually.

Write criteria as **behavioral anchors**: descriptions of observable behavior, not adjectives. "Names a receipt, invoice, or attachment" is an anchor; "well documented" is a mood. The anchor is what decides borderline cases, so write it for the borderline: the clear cases take care of themselves.

## Designing for independence

Questions are **independent by design**: answers come back under the same keys you sent, and no answer ever becomes context for another question. S01E01 stated that property; this section is about the authoring discipline it asks of you.

The failure mode to watch for is the **compound question**: one question that embeds two judgments. Compare:

- `"is_urgent_and_should_escalate"`: urgent _and_ we should escalate. Two judgments fused into one number. When it returns 0.6, you cannot tell whether urgency or the escalation logic is the uncertain part, and no downstream threshold can either.
- `"is_urgent"` and `"should_escalate"` as separate questions: two probabilities, each readable on its own, each routable on its own.

The rule: **one judgment per question**. Split compound questions into their parts and combine the parts in your code, where the combination is visible and changeable. The genuine exceptions are questions whose meaning changes after another answer exists, and those become a second request that uses the first request's answers in its state. The vendor documents the second request as "the exception, not the rule"; design as if it were rarer than you think.

The payoff for keeping questions independent is mechanical: independent questions batch into one request, and the fan-out economics from S01E01 carry over unchanged. Output tokens are free (`usage.input_tokens` is the whole bill), so asking a question you might not need costs almost nothing.

## A complete request: one expense line

Everything above fits in one real request. The scenario: an expense system that triages each submitted line automatically. The state carries the line and the numbers it must fit; the questions cover the three types and the independence rule.

```
POST https://api.typesafe.ai/v1/systemone
Authorization: Bearer $TYPESAFE_API_KEY

{
  "state": "Expense EXP-2041. Merchant: Grand Hotel Prague.
            Amount: EUR 412.80. Claimed category: lodging.
            Memo: '2 nights, client workshop; receipt emailed
            to finance.' Employee plan: standard.
            Month total so far: EUR 1,890 of the 2,500 cap.",
  "model": "jev-latest",
  "questions": {
    "receipt_referenced": {
      "type": "noul",
      "instructions": "The memo references a receipt or invoice",
      "criteria": {
        "true":  "Names or clearly references a receipt, invoice, or attachment.",
        "false": "No receipt is referenced; the claim rests on the memo alone."
      }
    },
    "within_cap": {
      "type": "noul",
      "instructions": "The line fits under the remaining monthly cap",
      "criteria": {
        "true":  "Adding the amount to the month total stays at or under the cap.",
        "false": "Adding the amount to the month total exceeds the cap."
      }
    },
    "lane": {
      "type": "choice",
      "instructions": "Who should handle this expense line",
      "criteria": {
        "auto_approve":   "Within policy limits, routine business expense",
        "manager_review": "Over a limit or policy-adjacent; a manager signs off",
        "finance_audit":  "Possible duplicate, split billing, or odd merchant",
        "other":          "None of the above fits cleanly"
      }
    },
    "policy_risk": {
      "type": "score",
      "instructions": "Risk that this line violates expense policy",
      "criteria": [
        "Routine: well within policy, nothing unusual",
        "Minor: technically claimable, but worth a second look",
        "Serious: likely policy violation, needs a ruling",
        "Severe: clear violation; flag for finance"
      ]
    }
  }
}
```

The response comes back under the same keys:

```
{
  "answers": {
    "receipt_referenced": 0.86,
    "within_cap": 0.93,
    "lane": {
      "value": "auto_approve",
      "confidence": 0.71,
      "probabilities": {
        "auto_approve": 0.71, "manager_review": 0.21,
        "finance_audit": 0.05, "other": 0.03
      }
    },
    "policy_risk": {
      "value": 1.5,
      "confidence": 0.44,
      "probabilities": [0.08, 0.44, 0.41, 0.01]
    }
  },
  "usage": {"input_tokens": 412}
}
```

Read it like your code will. The two nouls are bare probabilities: a receipt is referenced (0.86) and the cap holds (0.93). The choice names `auto_approve`, but look at the distribution before trusting it: `manager_review` at 0.21 is not nothing. The score is the teaching case: `1.5`, torn between "minor" and "serious", with the probability mass split 0.44/0.41 across the two levels. An integer-only API would have collapsed that hesitation; here your code sees it and can route the line to a human, which is the right answer for a borderline expense.

And those typed values are exactly what S01E03 consumes: threshold routing takes probabilities like these and turns them into pass, review, and block decisions your code owns.

## Knowledge check

**How much of the state does each question in a request see?**

- Only the questions it is paired with
- **✓** Exactly the same full state text
- A summary the model builds
- Whatever the first question returned

Every question sees exactly the same full state text and nothing else. The state is the only channel in: facts a question needs must be present when the request is written.

**A choice menu does not contain the true answer. What is the fix?**

- Raise the confidence threshold
- **✓** Add an escape option like 'other' to the map
- Ask the same question again
- Switch the question type to score

A menu forces an answer even when the true answer is absent. Your schema creates the escape hatch: author an 'other' or 'none of the above' option so reality has somewhere to land.

**A score question returns 1.5 on your 0-3 level scale. What does that mean?**

- The model errored; retry the request
- **✓** The model is torn between levels 1 and 2
- The scale must be re-indexed
- Only integer scores are valid

Scores can land between levels when the model is torn; 1.5 means the probability mass sits across levels 1 and 2. Your code decides how to interpret the fractional position.

**Question B only makes sense once you know question A's answer. What do you do?**

- Put A's answer in the state of the same request
- Merge A and B into one compound question
- **✓** Make a second request whose state uses A's answer
- Lower the criteria thresholds until both fit

No question sees another's answer, and a compound question entangles two judgments into one number. A true dependency becomes a second request, which the vendor documents as the exception, not the rule.

**Who decides what labels the model can return?**

- The model, from its training
- **✓** You, as criteria authored at request time
- The vendor's fixed label catalog
- The first option in the map

Every label is one you authored: true/false definitions for nouls, option descriptions for choices, level descriptions for scores. The model never invents labels; it locates the state inside the answer space you defined.

## Further reading

- [_Primitives (TypeSafe AI docs)_](https://docs.typesafe.ai/primitives): the concrete API surface this lesson compresses: question types, criteria shapes, response fields. Keep it open while writing your first requests.
- [_Quickstart (TypeSafe AI docs)_](https://docs.typesafe.ai/introduction/quickstart): a minimal first request end to end; useful as a shape check against the worked example above.

> **🧭 Practice** Your turn. Open the TypeSafe Playground, paste one real item you routinely judge (an approval, a ticket, a submission) as the state, and build one request: two noul questions with behavioral anchors as their true/false criteria, one choice whose map includes an escape option, and one score with level descriptions you would defend. Compare the typed answers with the decision you would have made by hand. If an AI assistant is part of your workflow, Copy page in the header menu hands it this whole lesson as markdown, citation included. Ask it to quiz you on the knowledge check, or to critique the criteria you wrote in the practice exercise. The glossary is the canonical vocabulary when terms blur.

## See also

- [📖 Jev Glossary](/courses/jev-fundamentals/glossary/)
- [📐 S01E01: Jev answers instead of prose](/courses/jev-fundamentals/s01e01-jev-answers-instead-of-prose/)
- [🎯 S01E03: Confidence, calibration & composite scoring](/courses/jev-fundamentals/s01e03-jev-confidence-calibration-composite/)
---

Source: [Jev State and Typed Questions](https://peripatos.dev/courses/jev-fundamentals/s01e02-jev-state-and-typed-questions/), Peripatos — free to read; reuse by permission.
