Jev Makes Agent Decisions Almost Free
TypeSafe’s new model charges for what it reads, not what it decides, forcing product teams to split agent work by economics.
AI-written, human-edited, never fabricated. How this is made
The cheapest part of an agent may now be its judgment. TypeSafe launched Jev this week with a pricing model that charges $0.042 per million input tokens and nothing for output. That changes the product question. Instead of asking which general-purpose model should run the workflow, PMs can ask which parts of the workflow need language at all.
That distinction matters because many agent tasks are not writing tasks. They are decisions wearing a conversational costume: approve this refund, rank these leads, flag this message, choose the right policy, score the support ticket. Jev returns typed probabilistic decisions rather than prose. Its available forms are Choice, Score, and Noul, according to TypeSafe’s launch announcement. The output is designed to be consumed by software, not admired by a user.
TypeSafe calls Jev the first public model in its System One category. The company describes System One models as systems for structured judgment, with response times of 70 to 500 milliseconds and zero hallucinations on structured outputs. That last claim is TypeSafe’s, not an independently established fact. The useful product detail is less grand: the model does not need to explain a decision when the product only needs a decision.
The launch arrived on September 15. The pricing documentation was confirmed on September 19, putting the most important number in plain view just as the model began attracting developer attention. The number is $0.042 per million input tokens. One hundred million input tokens would cost $4.20. One billion would cost $42, before the surrounding application costs. Output is free.
That is a peculiar reversal of the usual model bill. A generative model spends money producing language, even when the application immediately throws most of that language away. Jev makes the discarded part free because its output is not a stream of tokens in the usual sense. It is a structured decision. If you ask it to classify an item, the economical answer is not a shorter explanation. It is no explanation.
The workflow gets split
This pushes a new architecture toward the product roadmap. A general model can handle the parts where the user needs synthesis, ambiguity, or an answer in natural language. A decision model can handle the repeated gates around that answer. The agent might use a frontier model to draft a response, then Jev to decide whether the response meets a policy, whether the ticket belongs in escalation, or whether the customer is eligible for an offer.
That is not an infrastructure tweak. It changes what PMs should count as a product feature. A workflow that previously made one expensive model call may become several calls with different economic jobs. The generative model writes. The specialized model sorts, scores, or routes. The user may experience one agent. The bill will reveal the segmentation.
Consider an inbox assistant. It could send every message to a large language model and ask for a label, a priority, a suggested response, and a reason. That is a convenient prompt. It is also a poor cost model if the only product action is moving the message into a queue. A cheaper design would use a decision model for routing, reserve generation for messages that actually need a reply, and keep the explanation step conditional. The interface can stay simple while the backend stops pretending every task is composition.
The early usage report makes the point more vividly than the price table. In a test published by Every, Jev judged 777 items across 37 documents in 0.7 seconds for a quarter-cent, or $0.0025. The test is not a benchmark of every production workload. It is a small demonstration with an unusually clear result. High-volume classification can be priced like a background utility rather than a premium interaction.
The interesting question is no longer whether an agent can decide, but which model should pay for the decision.
The practical consequence is segmentation. PMs should stop treating model choice as a single application-wide setting. The unit of choice is the decision inside the workflow. A team may use one model for extraction, another for classification, and a third for the final user-facing answer. That sounds obvious when written down. Most agent products still begin with one model and add prompts around it. Jev makes that shortcut harder to defend.
There is also a pricing implication. If decisions become cheap enough, teams can run more of them. A product can score every inbound item instead of only the items a user opens. It can apply several policy checks instead of one broad instruction. It can use a model to decide whether another model should be called. The cost of the decision layer stops being the obvious constraint.
That does not mean inference is free. Jev charges for every input token, so verbose context remains a bill. A team that sends a full customer record, conversation history, policy manual, and tool trace for each classification has not discovered free intelligence. It has discovered a low output fee. The product work is still in the context window: deciding what the model needs to see, what it can ignore, and how often the decision should be recomputed.
The other counterpoint is capability. A typed answer is not automatically a correct answer. TypeSafe’s zero-hallucination claim applies to structured outputs, but structured output can still encode a bad judgment. A Choice can be wrong. A Score can be poorly calibrated. Noul, whatever its usefulness in a given workflow, still needs a product definition and an escalation path. Nobody outside the lab knows yet how Jev behaves across the messy distribution of real production inputs.
That uncertainty is why the right response is not to replace every model call with Jev. It is to give the model a narrow job and measure it there. For a classification gate, that means tracking agreement with reviewed decisions, false positives, false negatives, latency, and the cost of the calls it prevents. The cheap input price makes those experiments affordable. It does not remove the need for evaluation.
Cheap decisions, different products
OpenJev, listed on the project’s site and discussed after launch, points to the same pressure from another direction. Developers are already treating Jev’s economics as something to inspect, reproduce, and build around, rather than as a curiosity in a pricing table. The model’s public arrival was enough to produce early technical interest within the week. That is a distribution signal, not proof of product-market fit, but it matters when the central advantage is composability.
For PMs, the immediate planning exercise is straightforward. Take an existing agent flow and mark every step that ends in a small set of actions: yes or no, route A or B, score from a defined range, escalate or continue. Those are candidates for a specialized decision layer. Then mark the steps where the user needs a useful explanation, a novel synthesis, or a piece of writing. Those still justify a generative model. The point is not to make the diagram more sophisticated. It is to stop paying a language tax on decisions that never needed language.
This could change packaging too. Products have often priced AI features around messages, seats, or broad usage tiers because model costs were difficult to separate from user activity. A decision model introduces another possible meter: volume of decisions. That may be useful internally even if it never appears on the pricing page. A PM can estimate the cost of screening one million records, then decide whether the feature should run continuously, on demand, or only after a user triggers it.
The savings can also be spent. If classification costs $4.20 for 100 million input tokens, a team might use more checks, add a second-stage review, or pre-filter work before sending it to a costly generative model. The product may become more proactive because the economics allow it. That is the less obvious consequence of a lower price: not just a cheaper existing workflow, but permission to design a different one.
The catch is that cheap decisions can create expensive product behavior. More automated routing means more opportunities to route a customer incorrectly. More policy checks can produce more friction if their thresholds are wrong. More background inference can turn a quiet workflow into a system that is difficult to explain when something goes wrong. Cost optimization is not the same as product judgment. It merely makes judgment cheap enough to multiply.
Jev’s launch therefore puts a price on an old design habit. We have been using generative models as universal adapters because one model was easier to integrate than several. Now the specialized alternative has a number attached to it: $0.042 per million input tokens, with output at zero. The number is small enough that “just send it to the big model” is no longer a neutral default.
The next generation of agent roadmaps will likely separate the writer from the judge. The writer will remain expensive because users value its words. The judge will become cheap because the product only needs its decision. The teams that notice this first will not necessarily build smarter agents. They will build agents whose expensive intelligence is reserved for the moments when a user can actually see it.