AI Engineering
Jev and the Decision Layer: What a System One Model Changes in an Agent
A close reading of TypeSafe AI's Jev: state and typed questions in, calibrated Choice, Score, and Noul answers out. What is documented, what is vendor claim, where it fits in an agent, where it does not, and why it complements LLMs rather than replacing them.
- Author
- Akshay Shinde
- Published
- October 8, 2026
- Read
- 19 min read
- Topics
- AI Agents, Decision Models, Reliability, AI Infrastructure
AI Engineering / AI Agents / Decision Models / Reliability / AI Infrastructure
Jev and the Decision Layer: What a System One Model Changes in an Agent
A close reading of TypeSafe AI's Jev: state and typed questions in, calibrated Choice, Score, and Noul answers out. What is documented, what is vendor claim, where it fits in an agent, where it does not, and why it complements LLMs rather than replacing them.
On this page
- What actually ships
- The design pressure: atomic questions
- Confidence, precisely
- "Can't hallucinate", read carefully
- Where it breaks, by its makers' account
- The claims, and how much weight they bear
- What the docs say about agents, which is the interesting part
- Where it slots into an agent
- Where it should not be used
- Why "Jev replaces LLMs" misses the point
- What this means for reliable agents
A lot of production LLM code contains a ritual. Build a prompt that ends with "respond with only one of: billing, technical, sales". Call the model. Strip whitespace, lowercase the result, hope it did not add a sentence of explanation, fall back to a default when it did. Sometimes there is a retry. Sometimes there is a JSON schema and a structured-output mode, which helps with the shape of the answer and says little about whether the model was sure.
That ritual exists because a generative model's native output is a string, and most of what software wants from a model at a decision point is not a string. It is a branch.
Jev, released by TypeSafe AI in early access on 15 September 2026, starts from that mismatch. It does not generate text. You send it a state and a set of typed questions, and it returns typed answers with probabilities. TypeSafe calls it the first "System One model", after Kahneman's fast, intuitive System 1. Simon Willison, agreeing with Maggie Appleton, prefers "decision model", a name that describes what the thing does.
This article reads Jev from the documentation outward. The goal is to separate what TypeSafe documents, what it claims, and what follows as an engineering interpretation, and then to work out where a model like this belongs in an agent.
What actually ships
Everything below comes from the TypeSafe docs unless stated otherwise, as of jev-1.13.0.
There is one endpoint, POST /v1/systemone. A request has a state, a model, and a map of questions. The state is text: a string, a JSON object, or an array of text values. Images, audio, and video are not accepted. Each question has an ID you choose (the model never sees it), a type, instructions, and usually criteria.
{
"model": "jev-latest",
"state": {
"ticket": "I was charged twice for order A-104. Please refund the duplicate.",
"refund_policy": "Duplicate charges are eligible for a refund."
},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle `ticket`?",
"criteria": {
"billing": "Charges, invoices, refunds",
"orders": "Delivery, cancellation, returns",
"account": "Login, profile, security"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated does the customer appear in `ticket`?",
"criteria": ["Calm and factual", "Frustrated but civil", "Very angry"]
},
"refund_requested": {
"type": "noul",
"instructions": "Does `ticket` explicitly request a refund or credit?"
}
}
}The response carries an answer under each ID. A Choice returns choice, a probabilities map over every option, and confidence. A Score returns score, a position along the levels that can fall between them, plus probabilities, confidence, and a legend. A Noul returns a single noul: the probability that the answer is yes. Nouls have no separate confidence field, because the probability already is the uncertainty.
A few constraints shape how you use it:
- A Choice accepts up to 255 options. A Score takes 2 to 10 ordered levels.
- Questions in one request are evaluated independently and in parallel against the same state. One answer never becomes context for another. If a later question really depends on an earlier answer, that is a second request, issued by your code.
jev-1.13has a 64k-token budget per request and 32k for the state plus the longest single question.- English is the primary training language. Other languages work, less well.
- There is no fine-tuning. The same weights serve every account. You adapt Jev to a domain through the state, the instructions, and the criteria.
jev-latestis an alias that moves when a new version ships. The models page recommends pinning a version once you have tuned thresholds against it, which is good advice for any model and essential for this one.
The training method is called RLCD, Reinforcement Learning for Calibrated Decisions. TypeSafe positions it alongside RLHF and RLVR as a third post-training objective, optimized for "epistemically honest probabilities" rather than human preference. That is TypeSafe's description. This article did not find an independent technical description or evaluation of the method to check it against.
The design pressure: atomic questions
The most repeated advice in the docs is to decompose. Ask for "a judgment a knowledgeable person makes in a second given the right context." Do not ask Jev to "analyze this message and determine the best course of action." Ask whether it requests a refund, whether it mentions an open order, how frustrated the sender is, and combine the answers in code.
The advice follows from the architecture. Because questions are independent and cheap to add, ten narrow questions cost roughly the same latency as one. Because each answer is a typed value with probabilities attached, combining them is a weighted sum or a rule you can read in a code review. The composite scoring pattern makes the point well: when priorities change, you change a coefficient, not a prompt.
The docs push it further with speculative fan-out: ask every question your code might need, including ones that only matter on some branches, and ignore the irrelevant answers. Triage a ticket and score bug severity in the same call, and discard severity if it turns out to be a billing question. The primitives page cites a cookbook where batching 13 questions into one call was 11.5× cheaper and 9.6× faster than 13 separate calls; the cookbook's own summary gives slightly different figures (12.2× and 10.0×). The exact multiple matters less than the direction: in this model, asking more questions in one request is almost free.
The hidden cost is that the question set becomes the program. Someone has to decide which atomic judgments add up to "should this ticket be escalated", write criteria that separate options cleanly, and maintain them as the product changes. That is real design work, closer to building a feature model for classical ML than to writing a prompt. Jev moves effort out of output parsing and into question design. That is usually a good trade, but the work is still there.
Confidence, precisely
Confidence is the feature that makes Jev architecturally interesting, so it is worth being exact about what the number is.
For a Choice with n options, confidence is how far the top probability sits above an even split, rescaled to 0–1:
confidence = (p_max − 1/n) / (1 − 1/n)It is 0 when every option is equally likely and 1 when all the probability is on one. Only the top probability counts, so (0.6, 0.3, 0.1) and (0.6, 0.2, 0.2) both score 0.4. Score confidence uses a different formula that accounts for ordering: probability split between adjacent levels lowers confidence less than probability split between opposite ends. For a Noul, the docs suggest |2p − 1| if you want a comparable number.
Two things follow. First, confidence is not a separate signal from the model; it is a statistic computed from the same distribution you already receive. The docs say so openly and suggest alternatives, such as the top probability or the ratio of the top two, that may suit a particular decision better. Second, the calibration claim is about the probabilities, and it describes many predictions rather than one: outcomes assigned a probability of 0.8 should turn out correct about 80% of the time as a group, and any single one can be wrong. The confidence field is a rescaling of the top probability, not a probability itself. A confidence of 0.8 does not mean an 80% chance of being right; on a three-option Choice it corresponds to a top probability of about 0.87. The docs keep these two ideas separate, and thresholds should too.
What makes calibrated confidence useful is that it lets code take different actions for the same answer. The docs' voice-banking example routes everything below a floor to a person, shows a balance at modest confidence, and for a transfer either acts or asks the user to confirm depending on a much higher threshold.
The thresholds in that example are illustrations. The docs say directly that the right values depend on your domain and data, and that you should start conservative and plot confidence against accuracy on your own labelled cases. In other words, Jev supplies an uncertainty signal, and deciding how much of it to trust at each threshold remains the application's job.
"Can't hallucinate", read carefully
The launch post says Jev "can't hallucinate" and that it never makes type errors. The second claim is true by construction. The output space is the set of options you defined, so the model cannot return "biling" or a paragraph of apology. The launch post's own nuance notes that its 0% figure is "not empirical", because schema matching is guaranteed.
The first claim needs narrowing. Jev cannot invent a value outside your options. It can still pick the wrong option from inside them, confidently or not. In ordinary usage, "hallucination" covers both: fabricating content, and asserting something false. Jev removes the first kind at the boundary. The second kind is what calibration is meant to make visible, not impossible.
TypeSafe's own cookbooks make this distinction better than the launch copy does. The structured data extraction cascade contains the line that "schema validation is necessary but not sufficient: it catches structural errors, never semantic ones." The self-consistency cookbook runs one borderline moderation post through eight Choice questions, fifteen times per condition. Jev repeats its plurality labels 90.8% of the time, within the 87.5% to 100% range of the LLM configurations tested, and its probabilities vary less than in five of the six LLM conditions. Labels still flipped on 2 of the 8 questions. The cookbook's remedy is a minimum top probability of 0.60, with an explicit uncertain outcome below it that goes to human review. With that gate, Jev's agreement rises to 99.2%, with 74.2% of answers labelled automatically. One post and fifteen repeats is a small experiment; it illustrates the mechanism more than it measures the model.
The fair reading is that Jev makes decisions well-typed and inspectable, and gives code a signal for when to distrust them. Correctness still comes from good questions, good state, and code that treats low-confidence answers as unknown.
Where it breaks, by its makers' account
TypeSafe publishes a jaggedness page for jev-1.13, last reviewed on 2 October 2026. For an engineer deciding whether to depend on the model, it is one of the more useful pages in the docs, and it is direct:
- Literal reading. It answers the question you wrote, not the one you meant. Negations and implied conditions are taken at face value.
- Math and numbers. It does not count reliably, does poorly with numeric representations such as hex colors, and Score positions should not be used to interpolate exact magnitudes.
- Dates. It reads dates as text, not ordered quantities. Extract the parts as Choices, then compare them in code.
- Indirection. Double negatives and multi-hop questions cost accuracy.
- Large, irrelevant state. Accuracy falls as unrelated content grows. Filter first.
- Adversarial content. State is not treated as hostile by default. Injected instructions or text that argues for its own classification "can move the answer."
- Contradictory criteria, and option order: it can lean toward the first option in a Choice.
- Generation. It is not trained to produce text, and chaining Choices to fake it is slow and poor.
Some of these are ordinary engineering hygiene (keep arithmetic in code). Two matter more for agent builders.
The adversarial point means Jev is not, by itself, a prompt-injection defense. A Jev guardrail that reads untrusted content is subject to the same pressure as any model reading untrusted content. It may be cheap enough to run on every message, and an attacker cannot make it write harmful text, because it writes no text. But a cleverly framed input can still shift a probability, so a guardrail built on Jev needs adversarial testing like any other.
The literal-reading point means question wording matters more than it first appears. When a Jev decision is wrong and you catch yourself explaining what you meant, the docs suggest that explanation is the missing half of the instruction. That is good advice, and it is also a maintenance burden that grows with the number of questions.
The claims, and how much weight they bear
Speed and cost are the headline numbers on TypeSafe's own pages. Here is what the primary sources say, with the caveats TypeSafe attaches. Where a caveat is this article's reading rather than TypeSafe's, it is marked.
| Claim | Source | Caveat |
|---|---|---|
| $0.042 per million input tokens; output tokens free | Models page, launch post | Launch post: “We can't prove it isn't subsidized.” |
| 70–500 ms end to end; “most queries complete in about 100 ms” | Launch post; How to build | Evals are “generally” run from the team's laptops on the US West Coast, where the service is based. |
| 193.6× faster, 444.6× cheaper | Launch post, workflow evals | TypeSafe expects these to be “on the higher end of real world gains.” Workflows were built by its own team. The reference answers are an average of two frontier LLMs. |
| Similar intelligence to LLMs on System One tasks | Launch post | The workflow evals compare against LLM reference answers rather than ground truth. Reading: the post does not tie this headline claim to one specific eval. |
| 0% type errors | Launch post | True by construction, not measured. |
| 100K tokens/s, 80 requests/s | Models page | Limits “are adjusting dynamically” and may change without notice. |
All of these figures come from TypeSafe, and none of the argument here depends on an independent benchmark of them. The pricing is public, and anyone with access can check latency from where they are. The intelligence comparison is the claim to treat most carefully, because it is defined relative to other models' answers on workflows TypeSafe chose. TypeSafe is open about all of this, which is to its credit, and the caveats belong next to the multipliers whenever they are quoted.
What the docs say about agents, which is the interesting part
The How to build guide contains a sentence that reframes everything else: "System One is TypeSafe's model for building AI-powered software, not agents. It does not generate code or choose its own next action." It contrasts three architectures: traditional software, LLM agents ("every loop introduces another opportunity to go off the rails"), and AI-powered software, where code owns control flow and the model makes narrow, constrained judgments. It recommends avoiding agent while loops when a workflow can express the same behavior.
So the vendor's own position is not "use Jev to build agents." It is closer to "use Jev to need fewer of them." That is a coherent stance on one of the problems that make agents hard: in an agent, the model makes open-ended decisions about control flow at every step, and each of those is a chance to go wrong. A decision model offers a way to keep model judgment at those points while code keeps the control flow.
In practice, agents and AI-powered software are not either/or. A real agent is a loop containing many small decisions, and most of them are not open-ended. Which specialist handles this? Is the retrieved passage relevant? Does the draft cite something the source does not say? Is this action risky enough to need a human? Each of those is a bounded question with an answer space you can write down. In an agent built entirely around an LLM, each goes through the same generative model, as text or a tool call.
That is where Jev fits: not as the agent's brain, but as the decision layer inside it.
Where it slots into an agent
The docs' patterns and cookbooks cover several of the decision points an agent has. In each case Jev answers a bounded question and code acts on the answer.
- Routing. Classify intent and complexity, then send the request to deterministic code, a specialist LLM, or a person (intent routing).
- Tool and skill selection. The skill-suggestion cookbook ranks 182 agent skills against a user turn, re-checks the top three with fuller descriptions, and is allowed to reject them all. TypeSafe reports this cut incorrect skill loads by more than half. The function-calling cookbook maps a request onto a function name and closed-set arguments, each answered as a Choice with its own confidence.
- Escalation between models. The extraction cascade uses a cheap LLM first, Jev Nouls per field to ask whether anything looks wrong, and a reasoning model only when a flag fires. This is verification-triggered escalation; the cookbook does not use Jev to choose a model upfront. (It also ran on
jev-1.12.) - Guardrails. One request per message, on the way in and on the way out, with a Noul per hazard and a Score for severity, thresholded in code to pass, review, block, or route (LLM guardrails).
- Verification. The citation-check cookbook first does a plain string match to catch fabricated quotes, then asks a Choice whether the surrounding section supports, contradicts, or says nothing about the claim, and sends anything under a confidence threshold to a person.
- Fallback. Most of these patterns have a branch for answers that are uncertain or flagged. In intent routing, guardrails, and citation checking it ends at a person; in the cascade, at a stronger model; in skill suggestion, at loading no skill.
The following is an illustrative design rather than a TypeSafe example: one turn of a support agent handling a refund request, combining those pieces.
Code builds the state: the ticket, the order record, the refund policy. One Jev request classifies the request and asks a set of screening questions. Code applies the gates. On the automatic path, an LLM drafts a reply and proposes a refund, which is the part that needs generation. A second Jev request checks the draft against the state. Code enforces limits and executes the refund through an idempotent call. Anything uncertain, flagged, or with an amount outside the automatic range goes to a person.
const screen = await systemOne({ state, questions: screeningQuestions });
const team = screen.answers.team;
if (team.confidence < THRESHOLDS.route) return toHuman('unclear request');
if (screen.answers.injection_attempt.noul > THRESHOLDS.injection)
return toHuman('suspicious');
if (team.choice !== 'billing') return handOff(team.choice);
// The LLM proposes; it does not execute.
const draft = await llm.draftReply({ ticket, order, policy });
const check = await systemOne({
state: { ticket, order, policy, draft: draft.text, proposed: draft.action },
questions: draftChecks, // grounded_in_order, within_policy, promises_beyond_policy
});
// Two checks where yes is good, one where yes is a problem.
const draftOk = ['grounded_in_order', 'within_policy'].every(
id => check.answers[id].noul >= THRESHOLDS.check
);
const overPromises =
check.answers.promises_beyond_policy.noul >= THRESHOLDS.flag;
const amount = order.charges.at(-1)?.amountUsd ?? 0; // code reads the amount, never the model
const amountOk = amount > 0 && amount <= LIMITS.autoRefundUsd;
if (!draftOk || overPromises || !amountOk)
return toHuman('review draft', { draft, check });
await refunds.issue({
orderId: order.id,
amount,
idempotencyKey: `refund:${order.id}`,
});
return send(draft.text);systemOne here is a thin wrapper over the documented endpoint. The thresholds and limits are placeholders for values tuned on real data. Each draft question means what it says, so a yes on promises_beyond_policy is read as a problem in code rather than inverted inside the question; the jaggedness page notes that a Noul whose true maps to no performs worse. Note also what the code does not ask Jev: the refund amount, which code reads from the order record. Numbers are on the jaggedness list, and an amount that already exists in a record should come from the record anyway.
Where it should not be used
The first, fourth, and fifth points come straight from the docs. The others combine the docs with an engineering reading of where the architecture stops being a good fit.
- Anything that needs generated output. Replies, code, summaries, plans, explanations. The docs are explicit: use a generative model.
- Open answer spaces. If you cannot enumerate the options, or they change per request in ways you cannot build in code, a Choice does not fit. The docs' extraction pattern works around this by having regexes or an LLM propose candidates and Jev pick among them, which works when candidates are easy to generate.
- Decisions that need a stated rationale. Jev returns probabilities, not reasons. If a regulator, a customer, or an incident review needs to know why, you need the reasoning to come from somewhere else, or from the decomposition itself (the atomic questions and the weights are the explanation).
- Multi-step reasoning in a single question. By design, these should become several questions or several requests. If a decision cannot be decomposed, it is probably a System 2 task and belongs with a reasoning model.
- Numbers, dates, and counts, which belong in code.
- Untested languages and domains. Accuracy is best in English. Calibration measured on one distribution does not transfer automatically to another.
There are also trade-offs that come from the product, not the model. Jev is in early access from a young company, with rate limits that can change without notice. It is a hosted black box with no fine-tuning path, so the model itself improves only with new releases. Your own data can still shape the state and questions, or train a downstream classical model on Jev's outputs, as the docs describe. Aliases move, so a version upgrade can shift a calibrated system's behavior unless you pin it. None of this is unusual for a new model API. All of it belongs in the decision to put one on a critical path.
Why "Jev replaces LLMs" misses the point
It is tempting to read Jev as a replacement for LLMs in production. TypeSafe's own docs do not say this. The coding-agents page explicitly says Jev is not a drop-in replacement for the model behind a coding agent, and the jaggedness page sends generation back to "other models."
What Jev can replace is a particular use of LLMs: as an expensive, slow, unparseable classifier. Much of the infrastructure around an agent looks like that. The router that decides which prompt to use. The judge that decides whether an answer is grounded. The safety check that runs on every message. The "is this task done?" check at the end of a loop. Many of these sit on a generative model because it was the most convenient general-purpose tool at hand: no training data, no separate model to host. If a decision model does those jobs faster, cheaper, and with a usable confidence signal, LLM usage gets narrower, concentrated where generation and extended reasoning are actually needed.
That suggests a division of labour:
- The LLM proposes, reasons, plans, and writes. It handles open-ended work.
- Jev answers bounded questions: which, how much, whether. It handles the decisions that have a known answer space.
- Code owns control flow, enforces rules, does arithmetic, and performs side effects.
- A person handles what the gates mark as uncertain or high-impact.
The division is useful, and it has exceptions. Some routing decisions are better made by an LLM because the categories are fuzzy and change often. Some are better made by plain code, because a keyword or a database field already answers them and no model is needed at all. Some decisions should go to a person regardless of confidence, because the cost of a confident mistake is too high to automate. And Jev's confidence is only as good as its calibration on your data, which you have to measure before you trust it.
What this means for reliable agents
Beyond speed and price, Jev carries an argument about where reliability comes from. An agent built entirely around a generative model makes many small decisions per run through the same model it uses to write prose. Even with structured outputs, those decisions usually arrive without a calibrated confidence signal, mixed in with everything else the model is doing.
A decision layer pulls those decisions apart. Each becomes a named question with a defined answer space, evaluated on purpose-built state, returning a number that code can threshold, log, and test against labelled cases. The model is no smarter for it, but the agent's behavior becomes easier to inspect, because there are fewer places where control is handed to a model without anyone deciding to hand it over.
Whether Jev specifically becomes that layer depends on things the docs cannot settle yet: independent evaluation, behavior under adversarial input, price stability, and how well calibration holds outside TypeSafe's own workflows. The architectural pattern does not depend on Jev. Bounded questions, typed answers, explicit confidence, and code that owns the consequences are a good way to build agents with any model.