

What happened when we tested a model built to make decisions (vs. generate text) and where that primitive fits inside a production AI system.
Jev is taking the AI world by storm, so we wanted to see what the fuss was about.
While reading TypeSafe founder Diogo Almeida's launch post, something clicked—and once you see it, it is hard to unsee.
For the past few years, the dominant post-training recipes have optimized general-purpose LLMs for one of two targets: responses that human raters prefer (RLHF), or outputs whose correctness a verifier can score (RLVR). Both have made models dramatically more useful. But both still sit on top of the same primitive: an autoregressive model generates a string, one token at a time.
Then we embedded those models deep inside software and asked the string generator to impersonate a decision API:
Return this enum. Score this message. Produce exactly this JSON. Do not explain.
We changed the prompt and constrained the output, but we never really revisited the underlying primitive.
TypeSafe's counterproposal is more fundamental than “better JSON.” Jev is trained with what the company calls Reinforcement Learning for Calibrated Decisions (RLCD): optimize directly for bounded decisions with explicit probabilities, rather than prose that looks good to a person or earns a verifier's reward. In their framing: unstructured state in; typed, probabilistic decisions out.
Diogo explained this mismatch with unusual clarity: we optimized language models to be good at talking to people, then asked them to behave like dependable software interfaces. His framing is what made the idea click for us.
The obvious pushback is: isn't this just a classifier? Yes—often it is. That is the point, not a rebuttal. If software needs classifiers, we should build and train models to be classifiers from first principles: declare the answer space, score the alternatives directly, and expose the uncertainty. Structured Outputs remain useful; constrained decoding can enforce the schema. But that still puts a decision-shaped interface around a model optimized to generate strings. It does not change what the model was trained to optimize.
That is the hard-to-unsee mismatch.
A human-facing model needs to generate language. Software often needs a model to make a decision. Those are different jobs.
As an AI-native company, we use LLMs across our product stack. Jev made us ask a slightly uncomfortable question: how often are we paying a writer to do the job of a decision function?
At Attentive, AI Journeys turns customer behavior—such as browsing a product or adding it to a cart—into personalized, triggered SMS and email experiences. Each message is created in real time using subscriber, product, and brand context; we have written more about how those messages are crafted, from those inputs through evaluation and regeneration.
Now zoom in on one step inside that system. Suppose a generative model writes a 1-1 personalized SMS message using information about a customer, a product, the journey that was triggered, and the brand:

That is one generation.
Before software can act on it, the system may need to answer several other questions:
One generated artifact can create three, four, five, or six downstream decisions. That does not mean every check belongs in another model call. Exact rules should stay code. Ambiguous semantic questions are the decision workload.

The generation gets the attention.
The decisions make it a product.
As agents add verification, routing, retries, and abstention, the decision workload can quickly outnumber the generation workload. Better agents inspect their work, compare alternatives, enforce constraints, estimate uncertainty, and decide when not to act.
Most of that is judgment, not generation.
The conventional approach is to use a general-purpose generative model for both jobs.
Give a frontier model a rubric. Ask it to judge the result. Require JSON. Validate the JSON. Retry or fall back when needed. This works, and modern Structured Outputs make it work much better.
It can also be overkill.
When an application needs a boolean, a bounded choice, or a score, paying a prose generator to compose an answer—and then discarding nearly all of its expressive capacity—is like hiring a novelist to operate a switchboard.
Jev approaches the problem differently. It is designed to return bounded decisions: binary probabilities, choices from an application-defined set, and scores on an application-defined scale.
Conceptually, instead of asking a model to generate this:
{
"answer": "The message is probably acceptable because it follows the brand..."
}the application gets something closer to:
{
"decision": true,
"probabilities": {
"true": 0.93,
"false": 0.07
}
}
JSON is only the wire format. Semantically, the output is a typed decision: a value from a declared answer space plus probabilities. The judgment remains probabilistic. The software interface is bounded.
For a successful response, the selected value is constrained to the answer space the application declared. The model can still choose the wrong value.
Structured output itself is not new. We—and many other teams—have used schemas, function calling, parsers, retries, and validators around LLMs for years. OpenAI introduced strict Structured Outputs in August 2024. Those tools solve an important interface problem: they make a generative model's response reliably parseable and schema-conforming. They do not make the model's judgment correct or deterministic.
Strong frontier models also follow prompt-requested JSON almost all of the time. GPT-4o did so on every completed call in this pilot, and general Luna adhered to its strict schema on every completed call. But “almost all of the time” is not a software contract when JSON is requested only through a prompt. When response shape must be guaranteed, strict Structured Outputs is the right baseline.
What feels different about Jev is not JSON. It is the abstraction. A bounded decision is the native primitive, rather than prose generation constrained into a JSON envelope. The application declares the answer space, and the response includes a typed selection plus probability information that software can use for thresholds, routing, fallbacks, or review. Those probabilities still need to be evaluated for calibration.

We ran Jev 1.13, OpenAI's gpt-6-luna through both its Decisions and Responses interfaces, and gpt-4o-2024-08-06 on historical AI Journeys messages. Each model saw the same minimized message body and promotion context. GPT-4o was the pinned production control used by four core AI Journeys AutoQA judges.
This was intentionally a production-incumbent comparison, not an artificial contest over who can emit prettier JSON. Jev and the Luna Decisions arm used their native Decisions APIs and returned typed answers with probability information. General Luna used the Responses API with strict Structured Outputs and self-reported confidence. GPT-4o used our existing production AutoQA path: Chat Completions, temperature zero, and prompt-requested JSON.
When we began this pilot, we did not yet have access to OpenAI's Decisions API. OpenAI released the beta with gpt-6-luna shortly afterward, so before publishing we reran the same benchmark through it—and through general Luna using the ordinary Responses API. That makes the broader point clearer, not weaker: this is not a claim that Jev is categorically better than GPT. It is a test of whether decision-optimized models and interfaces are better primitives for decision-shaped work.
We evaluated Jev, Luna Decisions, general Luna, and our existing AutoQA pipeline (based on GPT-4o) on three LLM-as-a-judge checks:
Our question was narrow:
On bounded software decisions, what happens to quality, latency, and cost when we compare decision-native interfaces with ordinary generative interfaces—including the same Luna model on both?
The four arms landed within 0.7 percentage points on aggregate historical-label agreement across the supported checks: general Luna at 96.9%, Luna Decisions at 96.8%, the GPT-4o production control at 96.6%, and Jev at 96.2%. On this limited use case, that is not evidence of a semantic-quality winner; the same-model Luna comparison differed by just 0.1 percentage points.
The operational result was even clearer:
¹ Latency was measured end to end across four execution paths and three provider routes, so the comparison includes model serving, client, network, and routing differences.
² Cost is normalized over the same benchmark workload, using the identical bundled five-question request shape for every arm.
tldr: this pilot did not establish a quality winner. The same Luna model landed at 96.9% through Responses with Structured Outputs and 96.8% through Decisions. But the interface choice changed operations dramatically: Luna Decisions was about 11.3× faster at the median and about 25% cheaper than general Luna on the same bounded workload. Jev remained the lowest-cost arm; GPT-4o remained the production control. When one generation call can create several downstream judgments, that cost and latency structure matters quickly.
Both a frontier LLM and a decision model can return a valid typed payload. Both can put a decimal next to a choice. Those capabilities should not be collapsed.
There are three separate properties:
Structured Outputs is a major improvement to the first layer. A schema can require an enum and a number between zero and one. It cannot make the enum correct or the number honest.
Jev and Luna Decisions return probability information as part of their native decision interfaces. General Luna also returned a valid decimal confidence inside every strict-schema response, but that number was self-reported by a generative model—not a native probability distribution over the declared choices. TypeSafe says RLCD trains specifically for calibrated decisions. That is a more relevant optimization target for software, but it is not magic. This run did not establish that either decision model's probabilities are calibrated for Attentive traffic—or that a threshold would safely catch its mistakes.
A typed response can be perfectly valid and confidently wrong. Calibration is a population-level property. It requires independently reviewed cases, enough failures to measure the failure classes, and calibration curves by question and traffic slice.
Our claim is deliberately narrower: on the three checks the supplied evidence could reasonably support, all four arms landed within 0.7 percentage points, and the same Luna model differed by only 0.1 points across its two interfaces. Luna Decisions was about 11.3× faster at the median and about 25% cheaper than general Luna. The decision interfaces' native probabilities are an architectural reason to keep testing, not a calibration or quality claim we have already proven.
The benchmark result is interesting only if the primitive fits into reliable software.
In AI Journeys, an AI-generated message is a candidate—not an output. Product, brand, and journey context flow into a generator. The candidate then enters a loop where validators and policy decide whether to accept it, try again, or fail safely.
context → generate candidate → validate → accept, retry, or failThe generator proposes. The surrounding system decides what happens next.

AI Journeys has two different QA planes. Synchronous validators run inside the candidate loop and enforce exact constraints that may trigger another attempt. Asynchronous AutoQA evaluates a sampled subset more broadly, measuring quality and surfacing problems without pretending that every judgment is an inline send gate.
That distinction matters:
The clean integration point is a stable internal contract:
judge(state, questions) → typed answers + probabilities
apply_policy(evidence, thresholds) → accept | retry | review | failA provider adapter can translate that contract into Jev, an OpenAI model using Structured Outputs, or a future decision model. The response is evidence, not authority. Versioned application code owns thresholds, retries, failure behavior, and every side effect.
This is also the safe deployment interface. A new decision model can begin in shadow inside asynchronous AutoQA, where it cannot change a customer message. We can compare its decisions and probability distributions with the incumbent judges and independently reviewed labels before considering an inline role.
Typed evidence becomes much more useful when the system preserves its lineage. For every judgment, we should be able to reconstruct:
A compact decision receipt might look like this:
{
"generation_request_id": "...",
"candidate_id": "...",
"judge_version": "naturalness-v3",
"model_snapshot": "...",
"probabilities": {"pass": 0.81, "fail": 0.19},
"policy_version": "aij-candidate-policy-v7",
"action": "review"
}
That receipt turns model quality into an observable system. We can replay stored evidence through a proposed policy without another model call. We can separate a model change from a rubric change. We can shadow a new judge and measure which downstream actions would have changed.
Without that lineage, “LLM quality” becomes a pile of anecdotes.
The first wave of applied AI focused on generation.
Write the message. Create the image. Summarize the document. Produce the code.
The next wave will be shaped by orchestration.
Inspect the message. Compare the candidates. Enforce the constraint. Select the tool. Estimate confidence. Decide whether to proceed.
A capable agent may use a few expensive creative calls surrounded by a much larger layer of fast, inexpensive, typed decisions.
The most expensive mistake in agent design may be using a writer every time software needs a verdict.
Jev got our attention because of the numbers.
The idea behind it may matter more.
