Most AI models are designed to talk to people, but software needs to make decisions.
When a developer uses a large language model (LLM) to route a ticket or score a lead, they face a mismatch because LLMs produce text for humans to read. To use that text in code, the developer must coerce the model into a specific format and then parse that text back into a data type. This process is fragile, as a single word of conversational filler can cause the parser to fail and the system to break.
Jev is a System One model, applying the psychological concept of fast and intuitive thinking to AI. It does not generate text, but instead evaluates typed questions against a state and returns structured results directly.
The Decision Contract
Jev treats a decision as a function call. You provide a state, which can be a string, a JSON object, or an array of text, and a set of questions. Each question has a type: Choice, Score, or Noul.
A Score question asks “which level” and returns a numeric score and a legend that defines what each level means. A developer might use this to measure customer frustration on a scale from 0 to 2, where 0 is calm and 2 is very angry.
A Noul question asks “is this true” and returns a single value from 0 to 1. This is used for binary decisions, such as whether a request is urgent.
{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "technical",
"confidence": 0.78,
"probabilities": { "technical": 0.85, "sales": 0.0, "billing": 0.15 }
},
"frustration": {
"type": "score",
"score": 1.0,
"confidence": 1.0,
"legend": { "0": "Calm, just stating facts", "1": "Frustrated but civil", "2": "Very angry, strong language" },
"probabilities": { "0": 0.0, "1": 1.0, "2": 0.0 }
},
"is_urgent": { "type": "noul", "noul": 1.0 }
},
"usage": { "input_tokens": 392, "output_tokens": 65 }
}These questions run in parallel. Because Jev evaluates every question in a request independently against the same state, adding more questions barely changes the response time. This allows developers to perform a “speculative fan-out,” where they ask many questions at once and let the code decide which answers are relevant.
Trusting the Probability
The most important part of a Jev answer is not the decision itself, but the probability distribution behind it.
Jev is trained using Reinforcement Learning for Calibrated Decisions (RLCD). This is a third post-training approach, distinct from RLHF (which creates chatbots) and RLVR (which creates reasoning models). RLCD optimizes the model for calibration. A model is calibrated if an answer assigned a probability of 0.8 is actually correct 80% of the time across many predictions.
Code can use these probabilities to manage risk through confidence-gated routing. For a low-risk action, such as tagging a ticket for later review, a developer might set a confidence threshold of 0.5. For a destructive operation, such as deleting a user account, they might require 0.9. If the confidence falls below the threshold, the system can route the request to a human agent or request more information from the user.
The Cost of Generation
Traditional LLMs are slow and expensive because they generate text one token at a time. Jev skips this step entirely.
It also reduces cost. Jev 1.13 is priced at $0.042 per million input tokens, and the output tokens are free. In contrast, generative LLMs charge for both input and output, with output tokens often costing five times more than input tokens. This efficiency turns AI from a slow collaborator into a fast, reliable component of a software pipeline.
The Jagged Edge of Decisions
Like all AI, Jev has failure modes. Because it is a decision model and not a reasoning model, it can be literal. It answers the question exactly as written, not the one the developer meant to ask.
Jev is not a calculator. It does not count reliably and cannot perform complex arithmetic. Developers must keep all math in the application code. Similarly, Jev reads dates as text rather than ordered quantities. To compare two dates, the model should extract the dates as text, and the code should perform the comparison.
The model also struggles with indirection. Double negatives and multi-hop reasoning can reduce accuracy. Additionally, a state full of irrelevant detail can act as a distractor. To maintain accuracy, developers should filter the state in code before sending it to the model.
Machine Native Intelligence
Jev represents a shift toward Machine Native Intelligence, which is AI designed for AI-to-AI and AI-to-software interactions. In these systems, the machine interface matters more than the chat interface.
The industry responded to this idea quickly, and other “System One” models appeared within days of Jev’s launch.
{
"model": "laya-rl-agent",
"answers": {
"department": {
"type": "choice",
"choice": "billing",
"probabilities": { "billing": 0.4811, "technical": 0.2871, "sales": 0.2318 },
"confidence": 0.045,
"answer_confidence": 0.4811,
"action": { "act_probability": 1.0 }
},
"frustration": {
"type": "score",
"score": 0.9887,
"legend": { "0": "Calm, just stating facts", "1": "Frustrated but civil", "2": "Very angry, strong language" },
"probabilities": { "0": 0.0902, "1": 0.8308, "2": 0.0789 },
"confidence": 0.4798,
"answer_confidence": 0.8308,
"action": { "act_probability": 1.0 }
},
"is_urgent": { "type": "noul", "noul": 0.7603, "confidence": 0.7603, "answer_confidence": 0.7603, "action": { "act_probability": 1.0 } }
},
"usage": { "input_tokens": 205, "output_tokens": 0 },
"routing": { "model": "english", "repo": "convaiinnovations/laya", "reason": "English Latin text" }
}The Benchmarking Gap
Measuring these models is difficult because Jev’s weights and training data are closed. There is no single, independent ruler to measure every model under the same protocol.
Some models perform better on in-distribution data, while others handle out-of-domain tasks more reliably. Without open weights for Jev, the community relies on third-party experiments and browser-based tasks to estimate performance. For example, some benchmarks show Kev completing more browser tasks than Jev, while Jev remains faster at reaching specific web pages.
The Fuzzy If
Jev is not a replacement for coding agents or chatbots, but a decision layer that sits between intent and action.
It acts as a “fuzzy if-statement.” Where hand-written logic is too brittle to handle the variety of human language, Jev provides a structured branch. It allows a system to classify a request, score its urgency, and route it to the right tool before a generative LLM ever writes a single word.
This pattern allows for composite scoring, where a complex judgment is broken into several atomic scores. The code then combines these scores with weights to make a final decision. By placing the decision layer between intent and action, developers can build systems that are faster, cheaper, and more observable than those relying on a single, monolithic chat model.