Typed decision models: what early tests show

Models that answer typed questions with a probability in one pass: what they are, what the first independent tests found, and how to test one.

Chapters · 08

On 15 September 2026 TypeSafe AI released Jev in early access. It calls Jev its first “System One Model”, a name taken from Daniel Kahneman’s fast, intuitive System 1 thinking [1]. Jev does not write text. You give it a piece of state (a ticket, an email, a JSON object) and a set of typed questions, and it returns a choice, a score or a yes/no drawn from options you defined, with a probability attached. Three days later Convai Innovations published Laya, an open-weight model built on the same idea [11] [13]. Since version 0.3.7 it also ships a server that accepts the same request shape as Jev [11].

Both are days old. The idea still deserves attention, because it describes what many production AI systems do on every request: make a small decision, such as routing a ticket, moderating a message or checking an agent’s tool call, and know when to leave it alone.

The short version: Jev is fast, though the gap is large only against LLMs left in their slowest default modes [15]. On calibration, which is the point of the product, the independent tests are mixed [14] [15]. Laya is near chance until it is fine-tuned [11]. The comparison that should decide adoption is against a classifier fine-tuned on your own data, and the only one we have found is Laya’s own, on a benchmark its checkpoint was fine-tuned on [11].

What a typed decision model is

TypeSafe describes Jev as “a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out” [1]. A request has two parts. The state is the input: a string, a JSON object or an array of text. Images, audio and video are not supported yet [2]. The questions use three primitives [2]:

  • Choice picks one option from a list you supply, such as which team should handle a ticket.
  • Score places the state on a rubric, for example customer frustration on a scale from 0 to 2.
  • Noul is TypeSafe’s name for a yes/no question: it asks whether a statement is true and returns a probability between 0 and 1.

The Register gave an example of what a three-department routing question might return [17]:

{"billing": 0.08, "technical": 0.85, "sales": 0.07}

It paired that answer with a confidence of 0.82. The probabilities are per option. The confidence is a separate summary that TypeSafe derives from the shape of the whole distribution, and it is the number you set thresholds on; a Noul returns only its probability [5].

Jev’s internals are unpublished. TypeSafe says it built “a new model architecture, parallel sampler for maximum efficiency, and training method we call Reinforcement Learning for Calibrated Decisions (RLCD)” [1]. TechCrunch reports that the company has said little about that architecture, that outside observers suspect it sits on top of an open-weight LLM, and that the former OpenAI researcher who started the company says Jev is trained only on synthetic data [16]. So Jev can be judged only by what it returns, which is what the rest of this post does.

Laya is the clearest public account of how such a model can work, because its model card spells it out [11]. The English checkpoint is ModernBERT-large, a bidirectional encoder (it reads the whole input at once), fully fine-tuned, with a small decision head on top: 421M parameters in total. A multilingual checkpoint built on mmBERT has 322M. Each option in a question is scored at its own [MASK] token, the placeholder slot a BERT-style model is trained to fill, and the scores are normalised across that question’s options. The options arrive with the request, so a new schema needs no retraining, and every question in a call is answered in a single forward pass. Training rewards the model with strictly proper scoring rules, which pay the most only when the reported probabilities are honest.

Laya, then, is a BERT-style classifier whose label set arrives with each request, trained with an objective that rewards honest probabilities. That places the category between two tools most teams already know. A fine-tuned classifier is fast and accurate, but it needs labelled data and knows only its fixed labels. A zero-shot classifier takes its labels at request time, and one served as the baseline in an independent test below [14]. What typed decision models add is calibration as a training objective, several typed questions answered in one call, and an interface built for software to branch on [1] [11].

How this differs from asking an LLM to classify

Any LLM will produce a label from a prompt and a JSON schema. Three things are different, and the first is narrower than it looks.

The format is guaranteed by construction. “By providing typed, structured values, Jev can avoid the parsing and validating that must be done to process text responses from LLMs,” as The Register put it [17]. An LLM gets most of that guarantee from constrained decoding, which limits its output to a JSON schema or a forced tool call. The independent benchmark below ran every LLM that way: seven of the eight returned malformed answers on at most 0.7% of decisions in any suite, one small model reached 6.2% on one suite, and Jev returned none [15]. The format advantage is real but small, so the differences that matter are the next two.

The work happens in one step. An LLM writes its answer token by token; a typed decision model scores every option of every question at once. TypeSafe credits this parallel design for the speed [1] [3]. The design pays most when one call asks several questions: Laya answers all of them in a single forward pass [11], and TypeSafe says adding questions “barely changes the response time” [9].

The same routing question, asked two ways. The LLM writes its answer one token at a time, then a parser checks it. The typed decision model scores every option in one pass and adds a confidence value. Values from The Register’s example [17].

The probability is the product. A typed decision model is trained so that its per-option probabilities match outcomes, which is the aim of both TypeSafe’s RLCD and Laya’s proper scoring rules [1] [11]. LLMs asked to state their confidence tend to be overconfident, which Xiong and colleagues found across several models [20]. The research is mixed: Tian and colleagues found verbalised confidence from RLHF-tuned models can be better calibrated than their token probabilities [21]. Where an API exposes token log-probabilities, an LLM’s probabilities over the labels can also be read directly, and Kadavath and colleagues found large models well calibrated on multiple-choice and true/false questions posed in the right format [22].

Type safety guarantees the shape of the answer and says nothing about whether the answer is right. TypeSafe says Jev “can’t hallucinate” [1]; The Register noted that structured output with probabilities “does not preclude the possibility of being incorrect” [17]. And because the caller fixes the options, a Choice always picks one of them, even when none fits. TypeSafe’s own documentation says “the Choice is relative, settling which option, while each Noul is absolute and can be low for all of them” [4]. The practical fix is to add a “none of these” option, or to put a separate Noul in front of the Choice as a gate.

Why calibrated probabilities matter in production

A model is calibrated when its probabilities match reality across many predictions. TypeSafe’s primer puts it plainly: “Outcomes assigned a probability of 0.2 should occur about 20% of the time. Outcomes assigned a probability of 0.8 should occur about 80% of the time” [7].

Each small decision a production system makes carries two questions: what is the answer, and is it safe to act on. The CTO of Earendil, quoted by TechCrunch, described the trade: “The user has to say, okay, if this only comes back with 50% probability, maybe this is a coin toss, and I disregard it. But if it’s 95%, sure, then I can do something with it” [16].

TypeSafe’s documentation recommends three bands: act automatically when confidence is high; proceed with caution in the middle, for example by asking a person to confirm; and below that, route to a person, ask for clarification or fall back to another system. It adds that “a confidence threshold is not one number”, since actions with worse consequences deserve a stricter bar [5].

The operating number that follows is coverage at an error budget, an idea from selective classification, where a model is allowed to abstain [23]. Choose the error rate you can accept, say 5%. Set the threshold on labelled data so that errors among the items acted on stay within that rate, then measure what share of traffic the model handles alone. Everything else goes to a person or a larger model. One limit applies throughout: calibration is a property of groups of predictions, and in TypeSafe’s words it “does not guarantee that an individual answer is correct” [2].

A confidence gate turns a confidence value into an action. The bar rises with the cost of a mistake, so a decision at 0.80 is acted on at tagging and held before any money moves. Set your own thresholds on labelled data, one per action.

What the first tests show

TypeSafe’s blog gives end-to-end response times of 70 to 500 ms. Its home page claims Jev is “193.6x faster”, a figure the blog says comes from its workflow evals, in which each LLM ran at its provider’s default reasoning setting, and TypeSafe expects these “are on the higher end of real world gains” [1] [10]. The same post says the evals were made by its own capability team, “so some bias could exist”, and that the reference answers are the average of GPT-6 Astra and Claude Fable 5.1 [1] [10]. Accuracy in those evals therefore means agreement with two frontier LLMs, where most buyers would assume human labels. TypeSafe argues that this reference, if anything, favours the OpenAI and Anthropic models in the comparison [1].

Two independent benchmarks with published protocols followed within days. decision-model-benchmark ran Jev against eight LLMs from five providers, each constrained to a JSON schema or a forced tool call, with thinking turned off where the provider allowed it. It declares no vendor sponsorship and publishes its raw logs [15]. Calibration error below is expected calibration error (ECE), the average gap between stated confidence and observed accuracy; lower is better and 0 is perfect.

Finding Jev The eight LLMs
Accuracy, 77-way banking intent 76.3% 70.9% to 81.3%
Accuracy, SMS spam 93.0% 66.1% to 94.9%
Median latency per request 264 to 276 ms 303 ms to 5.6 s
Most options accepted in one question 255 512 (the most tested)
Low confidence (0.5 or below) when no option is correct 49.7% 97.3% to 100% for seven; 64.7% for one
Answers changed by shuffling option order 13% 15% to 37%
Calibration error, banking intent 0.083 0.043 to 0.226
Calibration error, SMS spam 0.249 0.042 to 0.197
Calibration error, suite built to force uncertainty 0.246 0.039 to 0.122

The authors’ summary is that “no class wins on quality”. Jev was the fastest model measured, about 1.2 times faster than the fastest LLM set-up (an open-weight model on specialised inference hardware) and 10 to 16 times faster than models in thinking mode, and they conclude that “the vendor’s speedup claim holds only against LLMs left in their slowest default mode” [15]. Calibration is mixed: Jev sits inside the LLMs’ range on banking intent and above it on spam and on the forced-uncertainty suite. The 49.7% is the relative Choice described above at work.

jev-benchmarks, a pre-registered pilot on 300 held-out examples against an open zero-shot classifier, found Jev more accurate on two of three tasks, where it could handle 83% and 86% of items alone at a 5% error budget [14]. On the third, emotion labelling, Jev put zero probability on the true label for 16% of examples and could automate none of them at that budget; the open classifier managed 2%. The authors fitted those thresholds on the same small slice, so they describe this sample only [14].

Laya documents its own limits unusually well [11] [12]. Its English base checkpoint scores 0.362 on a benchmark of four typed-decision workflows without fine-tuning: above the 0.318 random baseline, and below the 0.461 you would get by always guessing the most common answer. The headline 0.766 belongs to a checkpoint fine-tuned on that benchmark’s own training split. As shipped, the probabilities are overconfident. Fitting one temperature (a single number that sharpens or softens every probability) per question type and option count moves mean ECE from 0.466 to 0.081, per its model card, measured before a temperature change in a later release [11].

Held-out results are weaker than in-training ones. On held-out toxic chat, moderation accuracy is 0.530, barely above chance on a balanced split, and the authors put it bluntly: “hand-picked examples work, real traffic does not” [12]. The English checkpoint scored 0.000 accuracy on Khmer, and its authors warn that it stays confident while wrong, which is why their router picks a checkpoint by script before the model runs [11] [12]. The confidence figures from that language sweep were measured before the same temperature change; the maker has since re-run it, and accuracy held while the confidence figures moved [12]. A confidence gate protects you only on inputs like the ones the model was trained and calibrated on, so check the language and shape of inputs before the gate.

Laya’s card also has a table that puts it ahead of Jev on most rows, while stating that the Jev figures are “third-party published, never measured here” [11]. Its Jev calibration figure is the 0.246 from the forced-uncertainty suite above, and its own figure is measured after the temperature refit. Each side’s table suits its maker, and neither is a like-for-like comparison.

Both products are young. Jev is eight days old at the time of writing. At launch, access was through a waitlist [1], and TechCrunch reports that the company “briefly lost the ability to serve users from its API because demand was so high” [16]. Its rate limits “can change without notice”, and the same weights serve every account, so it cannot be fine-tuned on your data. TypeSafe says “Jev is not trained on customer requests or responses”, and that English is where its accuracy is best [3].

Laya’s first release reached PyPI on 18 September 2026 [13]. Its weights and code are open under Apache 2.0, so you can inspect and run it yourself. It needs fine-tuning and a temperature fit on your task first. Its escalation head, action.act_probability, “carries no usable signal yet”, and the card advises gating on confidence instead. Its bundled server listens on every network interface (0.0.0.0) with no authentication unless LAYA_API_KEY is set [11], so set the key before the port is reachable.

The evidence so far is vendor evals, one 300-example independent pilot and one independent benchmark of five small suites.

Choosing between an LLM, a typed model and a fine-tuned classifier

TypeSafe’s list of failure modes for jev-1.13 covers literal reading, arithmetic, date comparisons, questions whose answer depends on first answering another question, and counting: “jev-1.13 does not count reliably” [4]. Its advice is to keep arithmetic, dates and counting in code [4], and to have each question ask one well-scoped thing, composing the answers in code [8].

With that in mind, each of the three options has a place.

Option When it fits
An LLM The output is language, code or an explanation someone will read, or the decision needs several steps of reasoning. Jev “does not generate text, write code, or hold a conversation” [6]. An LLM can also draft labels for training a classifier, with people checking a sample, and give a second opinion on a low-confidence case before it reaches a person.
A typed decision model The options are fixed per request, one call asks several questions, and you have no labelled data yet. Used zero-shot, Jev beat an open zero-shot classifier on two of three tasks in one pilot [14]. Laya fits only after fine-tuning, since its base checkpoints are near chance [11], and its authors advise keeping Choice questions under about 20 options [12].
A fine-tuned classifier You have labelled data and the label set is stable. A 2024 study by Bucher and Martini found small fine-tuned models consistently outperformed zero-shot prompted LLMs on text classification [24].

How we evaluate a decision model before trusting it

  1. Build a held-out set from your own traffic, labelled by people and never used for prompt tuning or threshold fitting. Laya’s results show why: spam, which was in its training mix, scores 0.993, against the held-out moderation result above [12].
  2. Measure discrimination and probability quality separately. Accuracy and macro-F1 (which counts rare labels as much as common ones) for the first. For the second, Brier score and log loss, which penalise confident wrong answers, and calibration error with a reliability diagram like the one below.
  3. Fit a temperature on a separate calibration split and report calibration before and after. A hosted model’s weights cannot be tuned, but its output probabilities can still be recalibrated on your split. Temperature scaling has a single parameter and is often enough [18]. Laya found overconfidence on some tasks and underconfidence on others [12], so check each question type.
  4. Report coverage at your error budget for each action separately, since each action has its own threshold [5].
  5. Probe the edges. Shuffle option order, include items with no correct option, ask a question and its negation, and plant instructions inside the state. TypeSafe’s own docs show a question and its negation on the same ticket returning 0.72 and 0.47, which sum to 1.19 [4]. They also note that injected content “can move the answer” [4], which matters most when the model is a guardrail.
  6. Measure latency on the real path, at the median and the 95th percentile. Laya’s 33 to 40 ms is one question on a T4 GPU, depending on the checkpoint; on CPU its card gives 193 to 464 ms [11]. Jev’s figures are hosted round trips.
  7. Plan for drift. Pin the model version, since TypeSafe warns that an alias moves when a new release ships [3]. Log every probability, and recheck calibration on fresh labelled samples on a schedule. Calibration fitted on past data degrades when the inputs shift [19].
One temperature, fitted on a calibration split, pulls the bins close to the diagonal. A shift in the inputs pulls them off again, so the fit needs rechecking on fresh labels. The bars sketch the pattern; the error figures are from Laya’s model card [11].

Where this fits in how we build

We send classification to small trained classifiers and keep LLMs for work that needs language. Typed decision models share that view: a decision your software branches on should come from a model trained to return a calibrated label. Our view is that the most dependable path is still a small model fine-tuned on your own labelled data, for the reasons in the table above, and the gap between Laya’s zero-shot and fine-tuned results points the same way. We have not found a comparison of Jev with a classifier fine-tuned on a buyer’s own data, so that is the test worth running. Since Jev’s weights cannot be tuned per account [3], the tuning moves into how you word the questions and criteria, and into the thresholds in your code.

In practice we fine-tune encoder classifiers on a client’s data, calibrate and threshold their outputs, gate each action by confidence, and send the uncertain tail to an LLM or a person. Jev and Laya are worth testing against that baseline, on a held-out set, before either replaces it. If you have a classification step that is slow, expensive or hard to trust, tell us about it.

Sources

  1. TypeSafe AI, “Introducing System One Models & Jev”, 15 September 2026.
  2. TypeSafe docs, “System One”.
  3. TypeSafe docs, “Models” (jev-1.13.0, rate limits, fine-tuning, data handling, version pinning).
  4. TypeSafe docs, “Jev 1.13 jaggedness”, last reviewed 17 September 2026.
  5. TypeSafe docs, “Confidence”.
  6. TypeSafe docs, “Jev with coding agents”.
  7. TypeSafe docs, “AI primer”.
  8. TypeSafe docs, “Introduction”.
  9. TypeSafe docs, “Primitives”.
  10. TypeSafe, workflow evals.
  11. Convai Innovations, Laya model card (v0.3.7).
  12. Laya, BENCHMARKS.md.
  13. Laya on PyPI, first release 18 September 2026.
  14. jev-benchmarks, pre-registered pilot, September 2026.
  15. decision-model-benchmark, README and v2 report of record, September 2026.
  16. Tim Fernholz, “A new kind of AI model from a ChatGPT inventor is thrilling developers”, TechCrunch, 18 September 2026.
  17. Thomas Claburn, “TypeSafe AI debuts model for machines that plays Doom”, The Register, 16 September 2026.
  18. Guo, Pleiss, Sun and Weinberger, “On Calibration of Modern Neural Networks”, ICML 2017. arXiv:1706.04599
  19. Ovadia et al., “Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift”, NeurIPS 2019. arXiv:1906.02530
  20. Xiong et al., “Can LLMs Express Their Uncertainty?”, ICLR 2024. arXiv:2306.13063
  21. Tian et al., “Just Ask for Calibration”, 2023. arXiv:2305.14975
  22. Kadavath et al., “Language Models (Mostly) Know What They Know”, 2022. arXiv:2207.05221
  23. Geifman and El-Yaniv, “Selective Classification for Deep Neural Networks”, NeurIPS 2017. arXiv:1705.08500
  24. Bucher and Martini, “Fine-Tuned ‘Small’ LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification”, 2024. arXiv:2406.08660

Build this with us.

Bring the problem to a call, and the repository too if there is one.

Book a callRequest a code review
Experience
Two years serving customers, from startups to enterprise teams.
Shipped
30+ production AI systems.