System One models: A decision layer for enterprise AI workflows

Enterprise AI workflows involve a series of decisions before a system produces a response or takes an action. Consider an agent handling a customer’s disputed charge. It must identify the nature of the request, obtain the relevant account and transaction information, determine which team should review the case, and recognize when the available evidence is insufficient. If the agent updates the support case with a summary of its findings or changes its priority, the application must also enforce its authorization and approval rules.
These steps require different kinds of computation. Application code can check transaction amounts, permissions, and policy thresholds precisely. A language model can interpret the customer’s message or draft a response. Between those tasks are smaller semantic judgments: whether the message requests a refund, which category best describes the issue, or whether a proposed response needs further review. Such judgments may occur repeatedly within one agent run and across thousands of cases.
A general-purpose large language model can perform these tasks, including through a structured-output interface. However, many applications need a bounded decision and an indication of uncertainty rather than generated text. System One models are designed around that requirement. Given information about a situation and a set of predefined questions, they return typed answers and probabilities that application code can evaluate alongside deterministic rules. TypeSafe AI’s Jev is the first model introduced under this designation.
This article examines how System One models work, how they differ from generative language models and conventional classifiers, and what their probabilities and output constraints do, and do not, establish. It also explores practical use cases, implementation approaches, and the role these models can play within AI agents and enterprise AI solutions that require reliable routing, review, and control over consequential actions.
- What is a System One model?
- The engineering need for System One models
- How a System One Request Works
- The three decision primitives: choice, score, and noul
- How System One decisions are produced and evaluated
- Where System One models fit in an AI agent architecture
- Practical use cases for System One models
- Evaluating and adopting system one models in enterprise workflows
- From Jev to an ecosystem of System One models
- How LeewayHertz utilizes System One models to build enterprise agentic workflows
What is a System One model?
A System One model is a decision-focused AI model that evaluates information supplied by an application and answers predefined questions about it. Its input consists of a state, representing the situation to be assessed, and questions that specify the permitted forms of the answers. The model returns a set of typed decisions and probabilities that software can use directly. It does not generate an open-ended response or an explanation of its reasoning. TypeSafe AI introduced the category with Jev, its first System One model.
The state may be a passage of text or a structured record containing related information. In an IT incident triage workflow, it could include an alert description, the affected service, recent error messages, and responder notes. The application could ask which team should investigate first, whether the report indicates customer impact, and how severe the disruption appears against a defined rubric. Those are separate judgments about the same supplied information. The possible categories and rating levels are defined with the questions; the model evaluates the supplied state and returns answers within those definitions. TypeSafe AI’s implementation provides three question types, Choice, Score, and Noul, which the article examines in detail later. [1]
The name System One draws on the distinction between fast, intuitive System 1 thinking and slower, deliberate System 2 thinking. The analogy conveys the intended role of a fast, focused judgment within a larger process. It does not mean that the model replicates human thought, and it should not be treated as a technical description of its architecture.
In an application, a System One model can serve as a probabilistic decision component. It supplies a semantic judgment, such as the likely category of a request, while application code handles exact calculations, applies business rules, and determines whether the result should trigger a workflow step or be sent for review. The model makes the judgment; the surrounding system defines what that judgment is allowed to change.
Build AI agents with System One models
Design and build agentic workflows that use System One models at key decision points, evaluate their judgments against business cases, and integrate results with enterprise systems.
The engineering need for System One models
Many operations within an AI solution require focused judgments about unstructured information while advancing a larger task. A system may need to classify a request, select a workflow, rate the relevance of retrieved material, or identify a case that needs review. System One models address the cost and complexity of using a text-generating model at each of these decision points. [2]
- High volumes of bounded judgments: These tasks require an understanding of language and context, but their possible answers are often known in advance. A request may need to be assigned to one of several queues; a document may need a rating against a defined rubric. The application needs a usable result to continue the workflow.
- Latency across multiple decision points: An LLM generates output tokens sequentially. Although a single classification may be brief, an agent can make several dependent model calls during one task. Their response times accumulate, and repeating the same pattern across many tasks increases inference cost.
- A consistent interface for application code: Prompting an LLM to return a particular format may still require validation and error handling. Supported schema-constrained LLM APIs can enforce valid fields and values, so structured output itself is not new. System One models make typed decisions and probabilities their primary output, giving application code defined values to inspect and combine. [3]
- Information about uncertainty: A selected answer does not reveal whether it was a clear result or a close choice between alternatives. Probabilities provide additional information for workflow design. After testing their reliability on representative cases, a team can use them to determine when the application proceeds, gathers more evidence, or requests human review.
- Integration with exact business logic: A decision model can assess the meaning of a message or record, while application code calculates amounts, checks permissions, and enforces policy conditions. This division makes it easier to test each judgment and the rules that use it.
Together, these requirements explain the value of a decision-focused model within a larger AI system. Its practical advantage depends on how accurately and efficiently it handles the specific judgments the workflow requires.
How a System One Request Works
A System One request has two main elements: the state, which contains the information to be assessed, and the questions, which define the judgments the application needs. Each question specifies an answer type and, where applicable, the permitted options. The model evaluates each question against the shared state and returns typed answers and probabilities that application code can use in the next workflow step. [4]
Consider a support agent handling a disputed charge. The customer’s message, the status of relevant transactions, and an approved support policy form the state. The application needs to identify the team that should review the case and whether the customer has explicitly requested a refund.
{
"model": "jev-latest",
"state": {
"message": "I was charged twice for order A-104. Can you help?",
"transactions": [
{"order_id": "A-104", "status": "captured"},
{"order_id": "A-104", "status": "captured"}
],
"policy": "Billing disputes are routed to the billing team for review."
},
"questions": {
"review_team": {
"type": "choice",
"instructions": "Which team should review this case first?",
"criteria": {
"billing": "Questions about charges, payments, or refunds",
"technical": "Problems using the application",
"other": "The case does not fit either team"
}
},
"refund_requested": {
"type": "noul",
"instructions": "Does the customer's message explicitly request a refund?"
}
}
}
The request illustrates the steps involved:
- Assemble the state: The application gathers the message and relevant records from their source systems. The state should contain enough context to answer the questions without burying the relevant facts in unrelated information.
- Define bounded questions:
review_teamis a Choice question with three permitted answers.refund_requestedis a Noul question that asks for the probability of a yes answer. The question wording matters: requesting help with a charge is not necessarily the same as explicitly requesting a refund. - Evaluate the shared state: Both questions are evaluated against the same supplied information. They are independent: the answer to
review_teamis not passed torefund_requestedas part of this request. - Receive typed results: A simplified, hypothetical excerpt of the response could look like this. The numbers below are examples:
{
"answers": {
"review_team": {
"type": "choice",
"choice": "billing",
"probabilities": {
"billing": 0.94,
"technical": 0.02,
"other": 0.04
}
},
"refund_requested": {
"type": "noul",
"noul": 0.18
}
}
}
Here, the selected team is billing. The Noul value indicates a low estimated probability that the customer explicitly requested a refund. A full Choice response also contains a separate confidence field, omitted here to keep the excerpt focused.
- Apply workflow rules in code: The application can use the team judgment to propose a route, subject to thresholds established from testing. It can check transaction details through exact logic, enforce permissions before the agent updates the case record, and send uncertain or exceptional cases to a specialist. The model’s answer does not itself verify a duplicate charge or authorize a change.
- Record the decision: The workflow should retain the relevant source records, questions and criteria used, model version, returned probabilities, rule applied, and resulting route or reviewer action. These details allow the team to investigate errors and evaluate performance over time.
Questions that need the same state can be included in one request. If a later judgment depends on an earlier answer, for example, the selected team determines which policy must be retrieved, the application first obtains that answer, retrieves the additional information, and makes a subsequent request using the updated state.
Accelerate AI Solutions Development
Build fully functional solutions from your high-value use cases, based on specific operational needs and enterprise context.
The three decision primitives: choice, score, and noul
A System One question uses one of three primitives according to the answer the application needs. Choice selects among defined alternatives, Score assesses a position on a described scale, and Noul estimates the probability that a yes/no proposition is true. Each primitive has a distinct response format, allowing application code to handle the result according to its meaning.
Choice: Selecting from defined alternatives
A Choice question asks the model to select one option from a set supplied by the application. Each option has a name and can have a description explaining what it covers. The response contains the selected option in choice, a probability for every option in probabilities, and a separate confidence field. The confidence value summarizes how concentrated the distribution is: a clear lead for one option produces higher confidence than a close split across several options. It does not independently verify that the selected option is correct.
For the disputed-charge case, the application might ask, “Which team should review this case first?” and define billing for charge and payment issues, technical for problems using the application, and other for cases that fit neither description. An illustrative result might select billing with probabilities of 0.94 for billing, 0.02 for technical, and 0.04 for other. The selected team gives the application a proposed route; the full distribution shows whether another route was also plausible.
Option descriptions matter when categories overlap. A message about a failed checkout, for instance, could concern either a payment or an application error. Descriptions should state the boundary between those categories. Where the listed options might not cover a case, an other option provides an explicit outcome; where the state may lack enough information, an insufficient_evidence option can serve a different purpose.
Score: Assessing an ordered rubric
A Score question asks where a case falls on an ordered set of described levels. The application defines the levels from lowest to highest. Jev returns a probability for each level, a legend mapping level numbers to their descriptions, a score, and a confidence value that summarizes how concentrated the level probabilities are. Descriptions such as “routine request with no stated deadline” and “service interruption affecting current work” give the model clearer criteria than labels such as “low” and “high.”
Suppose the support workflow assesses case urgency using three levels: 0 for a routine inquiry, 1 for an issue disrupting the customer’s work, and 2 for an issue requiring prompt specialist review. If an illustrative response assigns probabilities of 0.60 to level 1 and 0.40 to level 2, the returned score is 1.4: the probability-weighted average of the level numbers. It lies between the two levels because the assessment is divided between them. The application should read the distribution alongside the score, since different distributions can produce the same average.
This is a rubric-based assessment, not an exact numerical calculation. A Score question can assess how urgent a message appears from its context; application code should calculate a transaction amount or the time remaining before a policy deadline from the underlying records.
Noul: Evaluating a yes/no proposition
A Noul question evaluates a single proposition and returns its estimated probability of yes as a number from 0 to 1. A value near 0 favors no, a value near 1 favors yes, and a value near 0.5 indicates a close judgment between the two. Unlike Choice and Score, Noul has no separate confidence field: the one probability expresses the model’s judgment over the two outcomes.
For example, a supplier writes, “Our remittance account has changed. Please use the attached details for future payments.” The application could ask, “Does this message request a change to the supplier’s bank details?” An illustrative Noul value of 0.96 would mean the model assigns a 96% probability to yes for that question. It would not mean the new bank details have been verified or that the requested change should be made; those are separate checks for the application and its reviewers.
The distinction between Noul and Score depends on the question. “Does this message indicate an urgent issue?” asks for the probability of a defined yes/no judgment. “How urgent is this issue?” asks for a position on a scale and requires described Score levels. Choosing the primitive carefully gives the application an answer with a clear meaning before it applies its workflow rules.
A workflow can use all three primitives on the same case: Choice to identify its category, Score to assess its urgency, and Noul to test a specific escalation condition. Application code can then combine those judgments with verified records and business rules to determine the next step.
How System One decisions are produced and evaluated
When Jev supplies a judgment to an application, the selected answer is only part of the result. Organizations also need to know how efficiently a set of judgments can be produced, whether the accompanying probabilities reflect observed outcomes, and what happens when a valid answer misinterprets the evidence. Those questions determine how much work the model can safely take on within a production workflow.
Decision inference and parallel sampling
A generative LLM produces output through autoregressive decoding. It predicts each token using the input and the tokens it has already generated, so later output positions depend on earlier ones. TypeSafe describes Jev as using a new architecture and parallel sampler to produce probabilities for predefined decisions.
The application can submit several questions against one supplied state. Jev evaluates each independently: the result of one question is not passed to another as additional context. This permits an early group of questions whose answers may be useful later, even if the application ultimately uses only some of them. TypeSafe calls this speculative fan-out. The approach trades a small amount of additional question processing for the possibility of avoiding another request. It works only when those questions can already be answered from the available state; newly retrieved evidence requires a subsequent stage.
There are two distinct sources of potential time savings. First, one request avoids repeated network round trips and repeated submission of the same state. Second, Jev evaluates the questions in parallel, so adding independent questions has relatively little effect on response time. Neither effect makes extra questions free, and a comparable LLM implementation could also issue independent calls concurrently. The relevant test is the whole decision workload, measured using the same inputs, questions, and acceptable error rate. Developers should compare median and high-percentile latency, throughput at expected concurrency, and cost per completed case.
The benefit of parallel decision sampling depends on the shape of the workload. Short states with many independent questions leave more room to reduce decision time; long inputs and workflows that require retrieval between stages introduce work that parallel evaluation cannot remove. Performance should therefore be assessed using representative state sizes and the decision dependencies the application will encounter.
Calibration, probabilities, and confidence
Calibration concerns whether predicted probabilities correspond to observed frequencies. Among a sufficiently large group of comparable predictions near 80%, the predicted outcome should occur in approximately 80% of cases. This is different from accuracy: a model can usually rank the right category first while assigning probabilities that are consistently too high. Calibration describes groups of predictions, never a guarantee about an individual result.
TypeSafe calls its training objective Reinforcement Learning for Calibrated Decisions (RLCD). Its stated aim is to produce decisions and probabilities that reflect uncertainty. This differs in purpose from using human preferences to improve generated responses, as in RLHF, or verifiable results to train performance on checkable tasks, as in RLVR. RLCD defines what TypeSafe intends to optimize; the resulting probabilities still need to be evaluated against the outcomes that matter to the adopting organization.
The distinction matters when interpreting confidence. For Choice and Score, Jev derives the confidence field from the spread of the probabilities it returns. A concentrated distribution produces a more decisive confidence value; a divided distribution produces a less decisive one. Confidence is therefore a summary of the model’s own distribution, not a second estimate of its historical accuracy. Low confidence can point to several different issues: overlapping answer definitions, a question that combines too many dimensions, or insufficient information in the state. Each calls for a different remedy.
Developers should use the probability that corresponds to the actual decision condition. If a rule is concerned with one Choice category, examine and validate the probability assigned to that category. If a Score rule concerns a critical band of several levels, the combined probability of those levels may be more useful than the weighted-average score. Reducing a distribution to one selected label, confidence value, or average can hide information the workflow needs.
A practical evaluation should answer five questions:
- What is the reference outcome? Define labeling rules and have qualified reviewers resolve ambiguous cases. Include rare categories and inputs with missing information.
- Are probabilities calibrated? Compare predicted probability ranges with observed outcome rates. Inspect results by category and relevant input group, not only in aggregate.
- What happens at each proposed threshold? Measure false positives, false negatives, the share handled automatically, and the volume referred for review. This exposes the tradeoff between automation coverage and errors among cases handled automatically.
- Do the error costs justify the rule? A missed case and an unnecessary referral may have different consequences. Thresholds should reflect those consequences and the team’s capacity to review uncertain cases.
- Does performance persist? Hold back a test set while refining questions and thresholds. After release, retain model and question versions, predictions, reviewer decisions, and eventual outcomes so that changes in language, case mix, or policy can be detected.
The results that matter most are those measured on the workflow where the model will be used. Cases with independently verified outcomes allow a team to determine its error rates, check whether predicted probabilities hold in practice, and choose thresholds that balance automatic handling with review.
Type safety and decision correctness
A bounded decision interface restricts results to the types and options defined for each question. If three categories are permitted, an unlisted fourth category cannot appear. This prevents unexpected output values, but a permitted category can still be the wrong judgment. Claims that a model “cannot hallucinate” should be interpreted in terms of this output constraint, not as a guarantee of factual accuracy.
A constrained answer space places more responsibility on question design. Options must cover the outcomes a workflow can encounter and distinguish neighboring categories clearly. TypeSafe’s Choice guidance recommends describing what an option includes and, where categories are easily confused, what belongs to another option instead. An other outcome can represent a case outside the taxonomy; an insufficient_evidence outcome can represent a case the supplied state does not support classifying. These outcomes are useful only if testing shows that the model selects them appropriately.
Even with well-designed options, the model can select a valid category that conflicts with the evidence or assign high probability to that selection. The application therefore has to establish four properties through separate mechanisms:
- Output validity: The model interface restricts the form and permitted values.
- Judgment accuracy: Labeled cases show whether the selected answer matches the evidence.
- Probability reliability: Calibration and threshold tests show whether reported probabilities support the proposed handling rule.
- Action authority: Permissions, exact business rules, and required approvals determine whether the next workflow step may occur.
A typed answer makes integration predictable. Its operational value comes from pairing that interface with measured decision quality and an application rule that specifies what the result is allowed to change.
Where System One models fit in an AI agent architecture
An agent does more than generate a response. It interprets a request, chooses what information to obtain, proposes actions, and evaluates what happened. Each transition can require a narrow judgment. A System One model can supply that judgment at a defined point in the agent loop, while the surrounding application determines how the result affects execution.
| Placement in the agent workflow | Judgment a System One model can make | How the agent uses the result |
|---|---|---|
| Before generation | Classify the user’s intent, identify the appropriate workflow, or assess which available model fits the task. | Route the request to a specialized workflow or select a model for the next step. An unclear result can take a default route or be reviewed before proceeding. |
| Before retrieval or tool use | Assess which of several available search paths or tools is relevant, based on their descriptions and the current request; identify whether the available state appears insufficient. | Choose a retrieval path or seek more information. If the next judgment depends on material retrieved at this stage, the agent evaluates it only after that material is available. |
| Before an action | Assess a proposed tool call against a defined risk criterion using the request, proposed operation, and relevant context. | Allow the application to pause, block, or refer the proposal for review. The application still checks permissions, exact policy conditions, and any required approval before executing the tool. |
| After generation | Assess a draft against a specific rubric, such as whether it addresses the request or includes a required qualification. | Revise or review a draft that falls below a tested threshold. If the question concerns support from source material, that material must be included in the state being evaluated. |
| After execution | Classify recorded events or agent traces to identify recurring failures, unusual paths, or cases worth investigating. | Prioritize monitoring and follow-up. The classification helps teams inspect activity; the underlying event records remain the evidence of what occurred. |
The model’s placement determines what it can judge. Before retrieval, it can select among described information sources, but it cannot assess the contents of a document that has not yet been retrieved. Before a tool call, it can assess the proposal; after execution, it can assess the recorded outcome. Designing questions around the evidence available at each point prevents an early judgment from being treated as though it had seen a later result.
LangChain’s Jev integration shows two concrete lifecycle placements. Its experimental model-routing middleware selects a model for an agent run based on the request. Its tool-call middleware assesses a proposed call before execution and can block a call classified as risky. The integration also allows custom middleware to evaluate agent state at other hooks and records classification traces alongside the agent’s activity. These are examples of where a decision component can be inserted, rather than a requirement to use a particular framework.
The decision and control boundary
A System One result is evidence for a workflow decision, not the decision’s governing authority. For a proposed action, the responsibilities remain separate:
- The model assesses a defined question, such as whether the proposed call appears risky or whether the available evidence is sufficient.
- Application code applies enforceable controls, including tool permissions, valid inputs, policy conditions, and tested probability thresholds.
- An approval mechanism obtains a person’s decision where the workflow requires one, before the action is executed.
This distinction matters in an agent harness. A risk classifier can flag or block a proposed call, but a low-risk classification does not grant permission to run it. The harness must enforce access rules and obtain human approval through a separate mechanism when required.
Each placement should also leave a trace of the information assessed, the question and criteria used, the returned result, the rule applied, and the eventual outcome. That record lets teams determine whether an agent took a poor path because of the model’s judgment, missing evidence, a threshold, or a control implemented elsewhere in the workflow.
Accelerate AI Solutions Development
Build fully functional solutions from your high-value use cases, based on specific operational needs and enterprise context.
Practical use cases for System One models
System One models can be used wherever a workflow repeatedly needs to interpret text or structured context and choose from a defined set of outcomes. Their value is clearest at a specific decision point: the application supplies the relevant state, asks a bounded question, and receives an answer with probabilities it can use for routing, prioritization, or review. The examples below are potential design patterns:
- Case intake and escalation: Given an incoming message, case history, reported impact, and descriptions of available teams, Choice can propose an owner with probabilities for each queue. Score can assess urgency against defined levels. The case system can use these results to propose an assignment and priority, while a triage specialist reviews ambiguous ownership or consequential changes.
- Document classification: From extracted text, record metadata, and a defined taxonomy, Choice can return a likely document type and probabilities for alternative types. A separate question can assess whether the text describes a semantic feature, such as an amendment to existing terms. The result selects a processing path; a specialist reviews uncertain classifications, while exact fields and reference numbers are checked in code.
- Invoice-exception triage: After a matching system identifies a discrepancy between an invoice, purchase order, and receipt, the model can assess the records and explanatory notes. Choice returns a proposed exception reason, such as an explained partial delivery, a disputed item, or an unresolved cause. Code calculates the numerical difference; accounts payable, procurement, or receiving verifies the explanation and resolves the case.
- Knowledge retrieval and evidence assessment: Given a question and passages retrieved from approved sources, Score can rate each passage against a relevance rubric. Noul can estimate whether the supplied passages appear sufficient to answer the specific question. The application can prioritize the material or retrieve more; source accuracy and support for the eventual answer still need to be checked.
- Policy-exception screening: Given a request, the applicable policy text, and relevant case details, Noul can estimate whether the request appears to depart from a stated policy condition. Choice can propose an exception category where several types are defined. The result routes the case to a policy owner; exact eligibility rules and exception approvals remain outside the model.
- Quality-issue triage: From a written defect report, product details, and defined issue categories, Choice can propose a category while Score assesses the reported impact. These outputs help organize investigation queues. Quality staff verify the affected scope and determine any containment, corrective action, or disposition from the underlying evidence.
- Agent tool-call screening: Given the user’s request, a proposed tool call and its arguments, and the current workflow context, Noul can estimate whether the call appears outside the requested task or warrants review. The agent harness can pause or refer the call. Application code still enforces tool permissions, argument restrictions, and required approvals before execution.
- Generated-output review: A draft response, the original task, a defined rubric, and relevant approved sources form the input. Score can assess whether the draft addresses the task, while Noul can test a specific condition, such as whether a required qualification appears to be missing. The result can prompt revision or human review; source-dependent and consequential claims still require verification.
For enterprise adoption, the useful unit is the individual judgment, not an entire department or agent. Each use case needs an identifiable source of evidence, answer definitions that cover expected and exceptional cases, an owner for uncertain results, and a clear next step. Those details make it possible to test whether the model improves the workflow while keeping subsequent actions under the appropriate controls.
Evaluating and adopting system one models in enterprise workflows
The starting point for enterprise adoption is a specific decision that slows down or complicates an existing workflow. A System One model is most useful when the decision requires interpreting context, the possible outcomes can be defined in advance, and the result has a clear role in the next step. Evaluation should establish whether the model improves that decision and whether the improvement carries through to the completed workflow.
Select and design the initial use case
- Identify a decision within the organization’s workflow. Examine where the current process requires a judgment based on text or other supplied context. Document how often it occurs, what information is available at that point, how the decision is made today, and what happens when the judgment is uncertain or incorrect. This establishes whether that specific decision is a suitable candidate for evaluation.
- Define the decision contract. Specify the source records that form the state, the precise question being asked, the permitted answers, and who owns the outcome. Include a route for cases that do not fit the categories or lack sufficient evidence. This makes the proposed model calls testable against the work people currently perform.
- Set acceptance criteria before testing. Decide which errors matter most, the maximum acceptable review volume, the required response time, and the cost target per completed case. For example, an exception-routing pilot might require fewer incorrect assignments without increasing the specialist queue beyond its available capacity.
Test against the existing workflow
- Build a representative case set. Use historical records with outcomes reviewed by the people responsible for the process. Include routine cases, rare categories, incomplete records, and examples on which reviewers disagree. Keep some cases aside for final testing so the evaluation does not simply reward question wording tailored to known examples.
- Compare practical alternatives. Run the same cases through the current rules or manual process and, where appropriate, a task-specific classifier and an LLM with constrained outputs. Compare the decisions each produces and the work required to obtain them. The best choice may differ by decision: exact eligibility checks belong in code, while a stable classification problem with extensive labeled data may justify a dedicated classifier.
- Evaluate outcomes and operating cost together. Measure incorrect assignments, missed cases, review referrals, response time under expected volume, and cost per completed case. Check whether the returned probabilities support the proposed routing thresholds on the organization’s cases. Include the time and cost of retrieving state, handling exceptions, and making any follow-up calls.
Introduce the model into production
- Run a shadow pilot. Let the model assess live cases while the existing process remains responsible for the outcome. Compare its proposed decisions with reviewer decisions and eventual results. This exposes differences between historical data and current work without changing case handling during the initial trial.
- Define the action and review paths. For each permitted result, specify what the application may do: propose a route, request more information, or refer the case to a person. Enforce permissions, policy conditions, and approvals in the application. A favorable model assessment can inform a workflow branch; it does not grant authority to take an action.
- Monitor and expand deliberately. Record the input sources, question definitions, model version, returned probabilities, selected path, reviewer changes, and final outcome. Reassess performance when policies, categories, source systems, or case volumes change. Pin the model version used for validated thresholds so an alias update does not silently change production decisions.
This approach gives an enterprise a practical adoption path: prove one bounded decision on its own cases, introduce it with a defined review boundary, and expand only where measured results show an improvement in the full workflow.
From Jev to an ecosystem of System One models
TypeSafe AI introduced System One models with Jev, but the approach is no longer tied to one provider or one way of building a decision model. Kev and Laya retain the central pattern of evaluating defined questions against supplied state, while making different choices about model architecture, customization, and deployment. Their value is to show how the category is developing beyond its first implementation.
- Jev — the hosted starting point. TypeSafe AI developed Jev as its first System One model and serves it through an API. It established a practical interface for sending state and typed questions to a decision model and receiving answers with probabilities that software can use. TypeSafe manages the model and serving infrastructure; the enterprise designs the questions, supplies the relevant context, and determines how the results affect its workflow.
- Kev — an open path to running and adapting the model. Jared Palmer developed Kev as a family of Jev-like models built on Qwen. Its published versions span different sizes, allowing teams to select a model according to their available hardware and performance requirements. Organizations can run Kev themselves and train it on labeled examples for their own decision tasks. Kev thus extends the approach from using a hosted decision service to operating and adapting an open model, with the associated responsibility for serving and maintaining it.
- Laya — a different model architecture and language strategy. Published by Convai Innovations, Laya uses encoder-based checkpoints: one based on ModernBERT for English and another based on mmBERT for multilingual input. Its router can select a checkpoint according to the supplied text. Laya demonstrates that a System One-style interface need not depend on Kev’s Qwen-based design; it can also be implemented with models built for different input and language requirements. Like Kev, it can be run on an organization’s own infrastructure.
Together, these models show the ecosystem expanding along three practical dimensions: managed versus self-hosted operation, general versus task-adapted models, and different approaches to handling context and language. That range gives enterprises more ways to fit decision inference into an existing architecture. It does not establish a universal best model: each implementation’s judgments, probabilities, latency, and maintenance demands must be assessed for the workflow in which it will be used.
How LeewayHertz utilizes System One models to build enterprise agentic workflows
LeewayHertz has experience designing and developing AI agents that work across enterprise data, applications, and human review steps. We apply that expertise to build agentic workflows using System One models to make focused decisions that guide subsequent agent actions, such as retrieving information, routing requests, or initiating human review. This gives an enterprise a practical way to introduce the model into a working process, rather than treating it as a standalone API.
Design the agent around a defined process
We work with the client’s teams to map the process the agent will support: where information enters, which judgments determine the next step, what the agent may prepare, and which outcomes require a person. From that design, we specify the state and questions for each System One decision, along with the business rules that govern its use.
Integrate System One decisions into agent execution
Our developers can build a decision component that calls a selected System One model and makes its typed results available to the agent’s orchestration layer. Depending on the workflow, those results can guide routing, determine whether to retrieve more evidence, or identify work that needs review. We can combine this component with LLMs for drafting and explanation and with application code for calculations and exact policy checks.
Connect the workflow to enterprise systems
We build the integrations that let the agent obtain approved records and return outcomes to the systems where work is managed. The agent can assemble context from relevant sources, retain references to the evidence used, and present proposed decisions in an interface that staff can review. These connections make the System One judgment useful within the client’s existing process.
Govern and improve the agent
LeewayHertz can configure the agent’s permissions, approval points, exception paths, and audit trail so a model result never becomes authority to act on its own. We then test the workflow against client cases and monitor decision quality, review volume, response time, and cost after deployment.
The result is an agentic workflow built around the client’s process, with System One models providing defined judgments at the steps where they add value. LeewayHertz can help an organization assess an existing agent, add a decision layer to it, or develop a new agentic solution with the integrations and controls required for production.
Endnote
An AI agent’s response or action is the visible endpoint of a longer decision process. Before reaching it, the agent may have selected a source, assessed whether evidence was sufficient, chosen a workflow path, or flagged a case for review. System One models make such judgments explicit through defined answers and probabilities that software can use at each transition.
This creates a clearer division of work within the application. A decision model assesses a focused question, a language model handles tasks that require generation, and code applies exact rules and permissions. For that design to work, the model must receive relevant evidence, its answer options must reflect the process, and its probabilities must be tested against the organization’s cases.
The value emerges in the completed workflow. Teams can trace which judgment influenced a path, examine how uncertainty was handled, and improve the component responsible for an error. System One models matter when that visibility helps enterprises route work more accurately, involve reviewers at the right points, and operate agents with greater consistency and control.
Ready to use System One models in an enterprise workflow? Partner with LeewayHertz to identify the right decision points and build AI agents that connect model judgments to your systems, business rules, and human review.
Start a conversation by filling the form
Once you let us know your requirement, our technical expert will schedule a call and discuss your idea in detail post sign of an NDA.
All information will be kept confidential.
FAQs
1. What is a System One model?
A System One model is an AI model built to make defined judgments about supplied information. An application provides the content to assess and typed questions that specify the possible answers. The model returns a selected category, a rubric-based score, or a yes/no probability, depending on the question. For Choice and Score questions, it also returns a probability distribution and a confidence value. These outputs let software use the judgment in a workflow.
2. How is a System One model different from an LLM?
A generative LLM produces text one token at a time. It can classify information, but it is also built for open-ended tasks such as drafting, explanation, coding, and reasoning. A System One model is built for decision inference: it evaluates questions with defined answer spaces and returns judgments with probabilities. That focus makes it suited to recurring classification, scoring, and yes/no decisions within an application.
3. What information does a System One request need?
A request needs a state containing the information to assess and questions defining the judgments required. Depending on the question, the application may also provide answer options or described scoring levels. The state should include the evidence needed at that point in the workflow, while the questions should be narrow enough for each result to have a clear use in application code.
4. Can a System One model answer several questions in one request?
Yes. The application can submit several questions about the same state and receive an answer to each. Those answers are independent: one question’s result does not become context for another question in the request. When a later judgment needs evidence that can only be retrieved after an earlier result, the application must gather that evidence and make a subsequent request.
5. What do the returned probabilities and confidence values mean?
A probability expresses the model’s assessment of a defined outcome. For Choice and Score results, a separate confidence value summarizes how strongly the probability distribution favors one option or level. It is not an independent measure of correctness. Before using a probability to determine when a workflow proceeds or requests review, teams should check how well it corresponds to observed outcomes in their own cases.
6. Does a constrained output mean the model cannot make an incorrect decision?
No. Constraining the output prevents the model from introducing an answer outside the permitted type or options. It can still choose the wrong permitted answer or assign it a high probability. Clear answer definitions and testing against labeled cases are necessary to assess judgment quality; the application must separately determine what action, if any, may follow.
7. What enterprise tasks are suitable for System One models?
They are suited to recurring decisions that require interpreting context and have answer spaces that can be defined in advance. Examples include request routing, document classification, exception triage, evidence relevance assessment, proposed tool-call screening, and review of generated content or agent traces. Each use case should have identifiable input evidence, a defined next step, and a way to evaluate whether the judgment improves the workflow.
8. Can a System One model control what an AI agent is allowed to do?
A System One model can inform a control point, such as assessing whether a proposed tool call warrants review. Its result may cause the agent workflow to pause, seek clarification, or refer the proposal to a person. The application or agent harness must still enforce permissions, argument restrictions, policy conditions, and required approvals before execution.
9. How should an enterprise choose between Jev, Kev, Laya, and other methods?
The enterprise should compare candidate methods on representative cases from the intended workflow. Relevant measures include decision accuracy, probability calibration, review volume, latency, cost, and deployment requirements. Jev is hosted, while Kev and Laya can be run on an organization’s own infrastructure. Existing rules, a task-specific classifier, and an LLM with structured outputs may also be useful comparison points.
10. How can LeewayHertz incorporate System One models into enterprise agentic workflows?
LeewayHertz can map a client’s process to identify decision points where a bounded judgment would help, evaluate suitable models against client cases, and integrate the selected approach with agent orchestration and enterprise applications. The workflow can then use model results alongside retrieved evidence, LLM-generated content, exact business rules, and human review, with each component assigned a defined role.
Insights
AI in portfolio management: Use cases, applications, benefits and development
AI is reshaping portfolio management by offering powerful tools that enhance investment strategies and decision-making.
How to build enterprise AI solutions for manufacturing?
AI in manufacturing leverages technologies like machine learning and deep learning neural networks to analyze vast data from various sources and facilitates improved decision-making by enhancing data analysis capabilities.
AI in legal businesses: Use cases, solution, benefits and implementation
AI reshapes legal firms by automating tasks, enhancing research capabilities, and providing data-driven insights, promising efficiency and client-centric outcomes.





