RAG & LLM Engineering  ·  BraivIQ AI Engineering Playbook

LLM Classification Under The Hood: How The New Decision Models Return Probabilities, And When Not To Use Them

LLM classification has a new shape. Instead of generating a label and hoping, the decision models released in the last four weeks score a closed set of answers and return probabilities. TypeSafe launched Jev on 15 September 2026, Cloudflare open-sourced Clef on 1 October and OpenAI put its Decisions API into beta on 6 October. The probabilities are only useful if they are calibrated on your own data, and an independent Red Hat benchmark published on 2 October found decision models no faster, cheaper or better than an LLM judge on its tasks. This deep dive explains how they work, how to measure them and when a fine-tuned classifier is still the better choice.

Published  ·  Updated  ·  14 min read  ·  By BraivIQ Engineering

Library card catalogue drawers labelled by letter, illustrating LLM classification into a closed set of categories

Key takeaways

  • A decision model scores every label in a closed set and returns probabilities, instead of generating text. Cloudflare's Clef scores the valid choices in parallel after a single prefill pass, and OpenAI's Decisions API returns predicate, choice and score answers with probabilities.
  • Vendors describe their probabilities as calibrated. That is a claim to test on your own labelled data, using expected calibration error and Brier score, before any threshold goes live.
  • Red Hat's benchmark of 2 October 2026, run from the UK, found decision models did not beat LLM-as-a-judge on speed, cost or quality, and that pre-trained classifiers such as DeBERTa remained very competitive.
  • A paper of 24 September 2026 showed natural-sounding context additions could flip 61.4% to 73.2% of correct decisions to a chosen wrong answer, so decision models should route work to people, not settle disputes.
  • Both code samples were run on Python 3.14.4 and 3.11.14 and pass mypy --strict. The Decisions API sample was type-checked against openai 3.27.0 and tested with real response types, not a live call.

15 Sep 2026 - TypeSafe launched Jev, a "System One" decision model  ·  1 Oct 2026 - Cloudflare open-sourced Clef under Apache 2.0  ·  6 Oct 2026 - OpenAI released the Decisions API in beta on gpt-6-luna  ·  61.4% to 73.2% - Correct decisions flipped by natural context additions in the JevOut paper (24 September 2026)

LLM classification is the workhorse of AI automation in financial firms. Which team gets this complaint, does this email mention a payment, is this client possibly vulnerable, is this alert worth an analyst's time. For two years the usual answer was to ask a chat model to write a label, parse it and hope. In the last four weeks a different shape went mainstream: decision models that score a closed set of answers and return probabilities. This deep dive is for senior engineers and CTOs deciding whether to use them. It covers how they work, how to check their numbers and when to choose something else.

The short version: decision models are a good fit for routing work to people, and a poor fit for decisions nobody checks. Their probabilities are only as good as your calibration, and an independent benchmark published on 2 October 2026 found they did not beat older approaches on the tasks it tested.

What are the decision models everyone was discussing in September?

On 15 September 2026 TypeSafe announced Jev, which it calls a System One model: "a new class of frontier models built to make fast, structured decisions that software can use directly". TypeSafe says Jev "outputs all probabilities in parallel instead of autoregressively generating by token", supports up to 255 possible outputs, is trained with what it calls Reinforcement Learning for Calibrated Decisions, and gives up free-text generation. It lists input at $0.042 per million tokens and end-to-end response times of 70 to 500 milliseconds measured from the authors' laptops on the US West Coast. It does not publish the model's size or licence, or details of the architecture.

On 1 October 2026 Cloudflare open-sourced Clef under Apache 2.0. Clef freezes Qwen3.8-27B, and Clef-flash freezes Qwen3.5-9B, then adds rank-256 low-rank adapters and a routing head. Cloudflare says "Clef uses Qwen for a prefill-only pass, then scores the valid schema choices in parallel", and that training used a Brier loss to refine probability calibration. It reports a BANKING77 macro-F1 of 94.20 for Clef, 90.93 for Clef-flash and 79.74 for Jev, on benchmarks "as defined by the Jev Decision Index", without sample sizes or test settings.

On 6 October 2026 OpenAI released the Decisions API in beta. Its changelog says it turns "text and images into typed answers 10x faster than the Responses API". The guide says gpt-6-luna is the only model, input costs $0.10 per million tokens and there are no output charges.

How does LLM classification work under the hood?

There are three broad ways to get a label out of a language model, and the new decision models are a refinement of the second and third.

  1. **Generate and parse.** Ask a chat model to write the label and parse the text. You get one sample, no probability, and occasional labels that are not on your list. Asking the model to state a confidence in words gives a number with no reliable link to accuracy.
  2. **Score each allowed answer.** Run the prompt once, then measure how likely the model finds each allowed label as a continuation and normalise across the set. That gives a real distribution over a closed set. It needs access to model probabilities, and the cost of scoring many labels can be shared by computing the prompt once, which is what Cloudflare describes for Clef.
  3. **Train a head for the job.** Put a classification head on top of an encoder such as DeBERTa, or on a frozen decoder as Clef does, and train it on labelled examples. One forward pass gives a softmax over the labels. This is the oldest approach and still the cheapest at scale.

OpenAI's Decisions API exposes the result rather than the mechanism. A predicate returns "an estimate from 0 to 1 that the condition is true". A choice returns one of your values with a probability for each and a separate confidence. A score rates against ordered levels and returns "a probability-weighted average" of their indices, "so it can fall between levels".

LLM classification architecture diagram: text and a closed label set go to a decision model that scores every label and returns probabilities, a calibration step fitted on labelled history adjusts them, a routing policy sets thresholds from the cost of errors and sends high-confidence routine items to a team queue and low-confidence or sensitive items to a person, and outcomes feed the labelled history used to recalibrate.
Decision models produce probabilities. Calibration and a routing policy decide what those probabilities are allowed to do. Tap to open full size.

Are the probabilities calibrated?

Calibrated means that across the items where the model says 80%, it is right about 80% of the time. Every vendor above claims calibration, and none publishes the calibration measurements you would need to rely on that for your own inputs. OpenAI's guide is honest about it: "Use labeled examples from your application to set thresholds for routing, filtering, or review".

One informal test shows why. On 7 October 2026 a developer posted on OpenAI's community forum that a predicate question about a coin weighted 70% to heads came out heads 70% of the time across 1,000 trials, while the same question as a choice came out heads 98% of the time. That is one person's test, not a benchmark, but it is exactly the kind of difference you will only find by measuring.

What does the independent evidence say about decision models?

Red Hat published the most useful independent comparison on 2 October 2026, updated on 5 October. Rob Geada, Mac Misiura and Shelton Cyril compared Jev, Laya and DiffusionGemma with LLM-as-a-judge guardrails and pre-trained classifiers on NeMo Guardrails benchmarks, running the evaluation from the United Kingdom. On prompt injection detection, Qwen3.6-35B as a judge scored 89.31% accuracy, DeBERTa 89.01% and Jev 86.35%. On content safety, Jev led with 86.20%. Their conclusion: "we did not find that decision models produced faster, cheaper, or higher-quality answers compared with LLM-as-a-judge", and "pre-trained predictive models remain extremely competitive".

Robustness is the other open question. A paper posted on 24 September 2026, "JevOut: Natural Context Can Flip Decision Models", added short pieces of background or procedural text that left the question and choices intact. Those additions redirected 61.4% to 73.2% of each tested system's correct decisions to a wrong answer chosen in advance. In a regulated firm, that is a strong argument for using decision models to route work to people rather than to make the final call.

What do AI automation code examples for decision models look like?

The first sample routes a client complaint with the Decisions API. It asks three questions in one call: which team, whether the client may be vulnerable and how urgent a response is. The model only chooses a queue. Anything uncertain, possibly vulnerable or possibly fraud goes to a person. It was written against openai 3.27.0, released on 9 October 2026, passes mypy --strict on Python 3.14.4 and 3.11.14, and was tested with real openai.types.Decision objects rather than a live call.

decisions_router.py
"""Route a client complaint to the right team with OpenAI's Decisions API (beta).

Python 3.11+, openai 3.27.0. Needs OPENAI_API_KEY. gpt-6-luna is the only model the
endpoint accepts. The model chooses a queue for a person. It never closes a complaint.
"""
from __future__ import annotations

from dataclasses import dataclass

from openai import OpenAI
from openai.types.decision import (
    AnswerAnswerResourceChoice,
    AnswerAnswerResourcePredicate,
    AnswerAnswerResourceScore,
    Decision,
)
from openai.types.decision_create_params import Question

MODEL = "gpt-6-luna"

QUESTIONS: list[Question] = [
    {
        "type": "choice", "name": "category",
        "instructions": "Which team should handle this client complaint?",
        "choices": [
            {"value": "fees_and_charges"},
            {"value": "payments_and_transfers"},
            {"value": "fraud_or_scam", "description": "a scam, an unrecognised payment or account takeover"},
            {"value": "service"},
            {"value": "other"},
        ],
    },
    {
        "type": "predicate", "name": "vulnerability",
        "instructions": "The client mentions circumstances that may make them vulnerable, such as illness, "
                        "bereavement, financial difficulty or a recent major life event.",
    },
    {
        "type": "score", "name": "urgency",
        "instructions": "How quickly does this need a response from a person?",
        "levels": [{"label": "routine"}, {"label": "within days"}, {"label": "same day"}, {"label": "immediately"}],
    },
]

# Set these from your own labelled history (see calibration.py). They are placeholders, not advice.
AUTO_ROUTE_MIN_CONFIDENCE = 0.90
VULNERABILITY_REVIEW_P = 0.20


@dataclass(frozen=True)
class Route:
    queue: str
    reason: str
    category: str | None = None
    confidence: float | None = None
    vulnerability_p: float | None = None
    urgency: float | None = None


def classify(client: OpenAI, complaint: str) -> Decision:
    return client.decisions.create(model=MODEL, input=complaint, questions=QUESTIONS)


def route(decision: Decision) -> Route:
    answers = {a.name: a for a in decision.answers}
    cat, vul, urg = answers.get("category"), answers.get("vulnerability"), answers.get("urgency")
    if not (isinstance(cat, AnswerAnswerResourceChoice) and isinstance(vul, AnswerAnswerResourcePredicate)
            and isinstance(urg, AnswerAnswerResourceScore)):
        return Route("triage_team", "a question was refused or missing")
    category, confidence, vul_p, urgency = str(cat.choice), cat.confidence, vul.probability, urg.score

    def to(queue: str, reason: str) -> Route:
        return Route(queue, reason, category, confidence, vul_p, urgency)

    if vul_p >= VULNERABILITY_REVIEW_P:
        return to("vulnerable_clients_team", "possible vulnerability, reviewed by a specialist")
    if category == "fraud_or_scam":
        return to("fraud_team", "possible fraud always goes to the fraud team")
    if confidence < AUTO_ROUTE_MIN_CONFIDENCE:
        return to("triage_team", f"category confidence {confidence:.2f} below threshold")
    return to(f"{category}_team", "routed on a calibrated threshold")


if __name__ == "__main__":
    text = "I was charged an exit fee twice when I moved my ISA last week. I've just lost my job so I need it back."
    print(route(classify(OpenAI(), text)))

The thresholds at the top are placeholders. The vulnerability check runs first and uses a deliberately low threshold, because missing a vulnerable client costs far more than an unnecessary specialist review. Refusals, which the API can return for any question, route to triage rather than failing.

How do you check calibration and set thresholds with LLM evaluation on your own data?

The second sample is the LLM evaluation step that should come before any threshold goes live. It measures expected calibration error and Brier score on a labelled sample, fits a temperature on one half, checks it on the other and picks the lowest confidence threshold that meets a precision target for automatic routing. It uses only the standard library, so it runs anywhere your labelled history lives.

calibration.py
"""Measure calibration on your labelled history, fix it with temperature scaling, and pick a
threshold for automatic routing that meets a precision target.

Python 3.11+, standard library only. The data below is synthetic sample data from a fixed seed,
built to behave like an overconfident classifier: 80% accurate, but 88% confident on average.
"""
from __future__ import annotations

import math
import random

Probs = list[list[float]]


def ece(probs: Probs, labels: list[int], bins: int = 10) -> float:
    """Expected calibration error: the gap between stated confidence and actual accuracy."""
    total, err = len(labels), 0.0
    for b in range(bins):
        lo, hi = b / bins, (b + 1) / bins
        idx = [i for i, p in enumerate(probs) if lo < max(p) <= hi or (b == 0 and max(p) == 0)]
        if idx:
            conf = sum(max(probs[i]) for i in idx) / len(idx)
            acc = sum(probs[i].index(max(probs[i])) == labels[i] for i in idx) / len(idx)
            err += len(idx) / total * abs(conf - acc)
    return err


def brier(probs: Probs, labels: list[int]) -> float:
    return sum(sum((p[k] - (k == y)) ** 2 for k in range(len(p))) for p, y in zip(probs, labels)) / len(labels)


def scale(probs: Probs, t: float) -> Probs:
    """Temperature scaling applied to probabilities: p ** (1/t), renormalised."""
    out = []
    for p in probs:
        w = [max(x, 1e-12) ** (1 / t) for x in p]
        s = sum(w)
        out.append([x / s for x in w])
    return out


def fit_temperature(probs: Probs, labels: list[int]) -> float:
    """Grid search for the temperature that minimises negative log-likelihood on held-out data."""
    def nll(t: float) -> float:
        return -sum(math.log(max(p[y], 1e-12)) for p, y in zip(scale(probs, t), labels)) / len(labels)
    return min((0.5 + 0.05 * i for i in range(91)), key=nll)


def pick_threshold(probs: Probs, labels: list[int], target_precision: float) -> tuple[float, float]:
    """Lowest confidence threshold whose auto-routed items meet the precision target. Returns (threshold, coverage)."""
    for t in sorted({round(max(p), 3) for p in probs}):
        chosen = [i for i, p in enumerate(probs) if max(p) >= t]
        correct = sum(probs[i].index(max(probs[i])) == labels[i] for i in chosen)
        if chosen and correct / len(chosen) >= target_precision:
            return t, len(chosen) / len(probs)
    return 1.0, 0.0


def synthetic(n: int, k: int, seed: int) -> tuple[Probs, list[int]]:
    rng = random.Random(seed)
    probs, labels = [], []
    for _ in range(n):
        y = rng.randrange(k)
        right = rng.random() < 0.8
        guess = y if right else rng.choice([c for c in range(k) if c != y])
        logits = [rng.gauss(0, 1) for _ in range(k)]
        logits[guess] += (5.0 if right else 3.2) + rng.gauss(0, 1.2)  # sure when right, a bit less sure when wrong
        z = [math.exp(v) for v in logits]
        probs.append([v / sum(z) for v in z])
        labels.append(y)
    return probs, labels


if __name__ == "__main__":
    probs, labels = synthetic(4000, k=5, seed=7)
    cal_p, cal_y, test_p, test_y = probs[:2000], labels[:2000], probs[2000:], labels[2000:]
    t = fit_temperature(cal_p, cal_y)
    fixed = scale(test_p, t)
    print(f"temperature {t:.2f}")
    print(f"ECE   before {ece(test_p, test_y):.3f}  after {ece(fixed, test_y):.3f}")
    print(f"Brier before {brier(test_p, test_y):.3f}  after {brier(fixed, test_y):.3f}")
    thr, cov_cal = pick_threshold(scale(cal_p, t), cal_y, target_precision=0.95)
    routed = [i for i, p in enumerate(fixed) if max(p) >= thr]
    precision = sum(fixed[i].index(max(fixed[i])) == test_y[i] for i in routed) / max(len(routed), 1)
    print(f"threshold {thr:.3f}: auto-routes {len(routed) / len(fixed):.0%} of test items at {precision:.1%} precision, "
          f"the other {1 - len(routed) / len(fixed):.0%} go to a person")

On its synthetic sample data, an overconfident classifier that is 80% accurate but 88% confident on average, the run prints a temperature of 1.50, expected calibration error falling from 0.079 to 0.033 and Brier score from 0.320 to 0.312. A threshold of 0.882 then routes 35% of test items automatically at 96.0% precision and sends the other 65% to a person. Your numbers will differ. The point is that you produce them, write them down and repeat the exercise whenever the model, the prompt or your inputs change.

What breaks in production?

  • **Thresholds copied from someone else.** A threshold only means something for one model, one prompt and one population of inputs. Set it from your labelled history and store the evaluation that justified it.
  • **Silent model changes.** A vendor update can shift calibration overnight. Pin model versions where you can, re-run the calibration check weekly on fresh labels and alert on drift in the routing mix.
  • **Labels that change meaning.** When a team splits or a product launches, the label set changes and old labels stop being comparable. Version the label set alongside the model.
  • **Mutually exclusive choices that are not.** A choice question picks one value. A complaint about fees that also reports a scam needs two predicates, not one choice.
  • **Adversarial input.** The JevOut results show context can steer a decision. Treat any input a client or third party wrote as untrusted, and never let a decision model's output close a case on its own.
  • **Data residency.** OpenAI lists Decisions API processing in the United States and in Europe (EEA and Switzerland), not the UK. UK GDPR permits transfers to the EEA under the UK's adequacy regulations, but your data protection impact assessment should record where the processing happens.

When should you not use a decision model, and what are the alternatives?

Use one when you have little labelled data, need several questions answered from the same input, or need image input. Choose something else when the task is stable and well-labelled, when you need the answer explained in words, or when the decision is final.

LLM classification options for a financial firm
OptionBest whenWatch out for
OpenAI Decisions API (beta)Several typed questions per input, little labelled dataBeta, one model, processing in the US or EEA
Clef, self-hosted (Apache 2.0)You need the model inside your own networkYou run the GPUs and own the calibration
Jev (TypeSafe)Lowest listed price per input tokenNo published size, licence or architecture
Fine-tuned DeBERTa or similarStable task with thousands of labelsRetraining when labels change
LLM-as-a-judge with structured outputYou also need a written reasonSlower, and probabilities need extra work
RulesThe criterion is exact, such as an amount or a codeBrittle on free text

For more on getting typed output out of generative models, see our guide to structured outputs and constrained decoding, and for wiring evaluations into release pipelines, agent evaluation in CI/CD. Our trading playbook on trade surveillance AI alert triage uses the same calibration step to order an analyst's queue. On the business side, our blog explains why measuring returns from day one separates the firms that get value from AI.

How can an AI Agency UK team help you put LLM classification into production?

Bring a sample of past items with the outcome your people chose, a few hundred is enough to start, and we will measure two or three options against it and recommend thresholds with the evidence attached. As an AI Agency UK team that builds signed-off AI for financial firms, we do this as the first week of a 14-day Proof Run, then put the winning option live with your people approving every outcome. The routing and triage agents we build are described on our services page.

Frequently asked questions

What is LLM classification?

It is using a large language model to put an input into one of a fixed set of categories, such as which team should handle a complaint or whether a message mentions a payment. The older approach asks the model to write the label. The newer approach scores every allowed label and returns a probability for each.

What is a decision model in AI?

In the sense used since September 2026, it is a model built to return a structured answer from a set you define, with probabilities, rather than free text. Jev from TypeSafe, Clef from Cloudflare and OpenAI's Decisions API on gpt-6-luna are examples. They trade away free-form generation for speed and typed output.

How do you get a reliable confidence score from LLM classification?

Use a method that returns probabilities over the labels rather than asking the model to state a confidence in words. Then measure calibration on a labelled sample from your own work, fit a temperature or similar correction if it is off, and set thresholds from the cost of each kind of error. Repeat when the model or your inputs change.

Is a decision model better than a fine-tuned classifier?

Not always. Red Hat's October 2026 benchmark found DeBERTa almost matched the best LLM judge on prompt injection detection at a fraction of the size. Decision models win when you have little labelled data, need several questions answered at once or need image input. A fine-tuned classifier wins when the task is stable and you have thousands of labelled examples.

References

  1. TypeSafe, "Introducing System One Models and Jev", 15 September 2026. https://typesafe.ai/blog/introducing-system-one-models-and-jev
  2. Cloudflare, "Clef: Open-weight decision models, and new RL fine-tuning platform", 1 October 2026. https://blog.cloudflare.com/clef-decision-models/
  3. OpenAI, "API changelog (Decisions API released in beta)", 6 October 2026. https://developers.openai.com/api/docs/changelog
  4. OpenAI, "Decisions (API guide)", Page checked 9 October 2026. https://developers.openai.com/api/docs/guides/decisions
  5. Red Hat Developer, "Benchmarking AI decision models against traditional guardrails", 2 October 2026, updated 5 October 2026. https://developers.redhat.com/articles/2026/10/02/benchmarking-ai-decision-models-against-traditional-guardrails
  6. arXiv, "JevOut: Natural Context Can Flip Decision Models (Xu et al.)", 24 September 2026. https://arxiv.org/abs/2609.30243
  7. OpenAI Developer Community, "Decisions API is now available in Public Beta (post 14, calibration test)", 7 October 2026. https://community.openai.com/t/decisions-api-is-now-available-in-public-beta/1403877
  8. Python Package Index, "openai 3.27.0", 9 October 2026. https://pypi.org/project/openai/3.27.0/