Future Product
Issue № 006 · 20/09/2026
Capability

TypeSafe Puts Typed Decisions Beside the LLM

Jev turns structured judgment into a separate product primitive, giving teams a reason to stop asking generative models to do every job.

AI-written, human-edited, never fabricated. How this is made

an agent workflow splitting a request into a generative writing path and a typed decision path

Your next agent architecture may need two models before it needs a bigger one. One model writes the email, explains the result, or talks to the customer. Another decides whether the email should be sent, which queue receives the ticket, or whether a proposed action clears a policy check. That second layer is the bet behind Jev, which TypeSafe launched on September 15 as the first public example of what it calls a System One model.

The important detail is what Jev does not return. It is a non-LLM model that returns typed probabilistic decisions, rather than text. TypeSafe describes three output types, Choice, Score, and Noul, and says responses arrive in 70 to 500 milliseconds. The company also claims zero hallucinations on structured outputs. That combination sounds narrow until you picture an agent stack full of tiny moments where a product needs a decision, not a paragraph.

Imagine an operations agent reviewing an incoming request. A generative model can read the request and propose an answer. But the product still needs a clean decision about whether the request belongs in billing, security, or support, perhaps with a score that tells the rest of the system how much confidence to require before acting. Today, teams often prompt an LLM to produce that judgment in prose, then parse the answer back into software. Jev starts at the other end: the output is designed to be consumed as a typed decision.

A model for the branch

This is the first useful conceptual shift. A Choice is not a mini essay that happens to contain a label. A Score is not a sentence saying “probably yes.” The output is meant to occupy a known place in a workflow. TypeSafe’s System One framing gives product teams a name for that layer, even if the company’s public launch post leaves plenty still to prove about how broadly the approach generalizes. TypeSafe’s introduction describes Jev as using RLCD training and a parallel sampler architecture, the machinery behind its attempt to make decisions fast, structured, and dependable.

That matters because parsing is not a product feature. It is a tax. When an LLM is asked to classify a request, the team has to define the prompt, constrain the format, handle malformed output, decide what uncertainty means, and test whether a confident-sounding answer is still wrong. A typed decision does not eliminate the underlying classification problem, but it removes one awkward translation step between model behavior and application logic. The agent can branch on the result directly instead of treating language as an API.

The phrase “zero hallucinations” needs a hard boundary around it. TypeSafe’s claim is about structured outputs, not universal correctness. Jev could return a perfectly valid Choice that is the wrong Choice. It could assign a Score that looks precise while the underlying judgment is poor. The launch material, as presented in the evidence available this week, does not give us an independent benchmark or a failure-rate comparison that would settle those questions. The honest reading is narrower and more useful: the model is designed not to invent extra text when the application asked for a typed result.

The useful question is not whether Jev can write, but whether your product should ask it to.

That distinction is easy to miss in the excitement around a new model category. Coverage from TechCrunch on September 18 describes the non-LLM design as attracting developers. I understand why. Developers have spent the last few years making language models behave like reluctant database fields, filling in schemas and then defending the boundary between an answer and an action. Jev proposes a model whose native job is closer to the decision node in a flowchart.

For product managers, the architectural consequence is more significant than the novelty of the label. If you treat the LLM as the agent, you tend to put more responsibility inside the prompt. The model reads the context, reasons about policy, chooses a tool, calls the tool, explains the call, and recovers when something goes wrong. A System One layer suggests separating those responsibilities. The generative model can handle interpretation and communication, while Jev handles selected decisions that need a stable type and an explicit handoff.

That separation changes where you draw the product boundary. Consider a support agent that drafts a refund response. The writing model might produce a warm, accurate explanation. A distinct decision model could determine whether the request is eligible for an automatic refund, return a Choice, and provide a Score for the system’s next step. A low score might route the case to a person. A high score might allow the workflow to continue. The example is architectural, not a claim that Jev has been proven on refund decisions. That proof is exactly what teams will need to establish in their own domain.

The decision contract

The new primitive is less about replacing a general model than about creating a contract between models and software. A typed result gives the product team something to specify before implementation: what are the allowed Choices, what does a Score mean, and what should happen when the model is uncertain? Those are product questions, not merely inference settings. They belong in discovery with the people who own risk, escalation, and customer experience.

This is where Jev could force better agent design even if it does not win every benchmark. A roadmap that says “add an agent” is too coarse. You now have to decide which parts of the experience require generation and which require judgment. The distinction can expose hidden assumptions. Is the agent allowed to choose an action, or only recommend one? Is a score a confidence estimate, a ranking signal, or a threshold input? Who owns the fallback when the output is valid but not useful? Typed decisions make those questions visible because they refuse to masquerade as conversation.

The latency range is part of that case. TypeSafe says Jev responds in 70 to 500 milliseconds, which is a wide band but still speaks to a different role from a model producing a long explanation. A decision layer can sit inside a live workflow without making every branch feel like a second chat turn. That does not mean every use case will feel instant, or that the published range tells us how it behaves under real production load. It does mean teams can evaluate decision quality and interaction timing separately instead of accepting one large model’s trade-offs for the whole experience.

The economics point in the same direction, though it is not the main story here. Jev launched at $0.042 per million input tokens with free output, according to the launch evidence. That makes frequent decision checks easier to imagine, but the more consequential change is conceptual: inference can be allocated by job. The writing model need not own every classification, and the decision model need not pretend to be a writer. Pricing may help teams make that split, but the split starts with capability.

Build the handoff

The practical work for PMs is to design the handoff rather than add another model to a diagram. Start with one decision that is currently hidden inside a prompt. Write down its valid outputs. Define what each type means to the user, the workflow, and the person who inherits an uncertain case. Then test the decision independently from the generated response. If the agent gives a beautiful explanation for a bad branch, you should be able to see that failure immediately.

There is a counterpoint, and it is substantial. A separate decision model adds another vendor, another evaluation surface, and another place for context to be lost. Some teams will prefer one capable LLM because the system is simpler to operate and easier to iterate. Jev’s public launch does not establish that System One models outperform a well-engineered LLM classifier, nor does it show how they behave across domains. Nobody outside the lab knows yet whether the clean separation survives messy, ambiguous inputs.

But “we can make one model do everything” has also been a seductive form of technical debt. It hides control logic in prompts, mixes explanation with authorization, and makes it difficult to tell whether an agent failed because it misunderstood a request or because it made the wrong decision. Jev’s arrival gives teams a concrete alternative to that habit. Not a universal replacement, and not proof that typed outputs are automatically trustworthy. A separate decision layer is simply easier to name, test, meter, and govern.

The strongest version of TypeSafe’s claim is therefore not that Jev is a better chatbot. It is that chatbots were the wrong abstraction for some of the most important steps inside an agent. If that holds, the product stack gets a new seam: generation on one side, typed judgment on the other. The teams that benefit will be the ones willing to put that seam in the architecture before a failure forces them to.

Sources