Jev: The AI That Decides Instead of Writes
Part of: AI Learning Series Here
Quick Links: Resources for Learning AI | Keep up with AI | List of AI Tools | Local AI | AI Agents | Future of Work
Subscribe to JorgeTechBits newsletter
Explore the Latest Token Prices
Disclaimer: I create this content entirely on my own time, and the views expressed here are mine alone (not my employer’s). Because I love leveraging new tech, I use AI tools like Gemini, ChatGPT, Claude, Perplexity and others as a “digital team” to help research and polish these articles so I can share the best possible insights with you!
Why an AI model that cannot write a sentence could become a useful building block for enterprise workflows—and what a sub-500-millisecond trading demonstration actually shows.
Most conversations about AI focus on what models can generate: articles, emails, software, images, and explanations.
But much of the work inside software does not require a paragraph. It requires a decision.
Which team should receive this request? Does this customer need immediate attention? Which document is relevant? Should an automated action proceed—or stop for human review?
In the last few weeks, Jev has become the talk of the town in AI circles—but it is not another text-generating LLM. Developed by TypeSafe AI, a startup founded by former OpenAI researcher Diogo Almeida alongside Erik Gafni and Sasha Sheng, Jev emerged in September 2026 with a different mission: making fast, structured decisions rather than writing responses. Almeida previously helped develop the instruction-following research behind ChatGPT; now, his team is pursuing an AI model that evaluates the information you provide and returns choices, scores, and probabilities for predefined answers that software can act on directly. TypeSafe calls it a “System One” model: less about composing an answer, more about deciding what happens next.
The simplest way to understand the distinction:
A generative LLM writes an answer. Jev evaluates the answers you have already defined.
That does not make Jev a replacement for ChatGPT, Claude, or other generative models. It makes it a different tool for a different part of the workflow.
Why decisions can be faster
A conventional generative LLM produces its output sequentially, token by token. Even when the result is structured JSON rather than prose, the model still generates that response.
Jev does not need to write the response. It returns probabilities for bounded choices and can evaluate multiple questions about the same input in parallel.
Consider a customer message:
“This is the third time I’ve contacted you. Our export still doesn’t work, and we cannot finish our monthly reporting.”
A workflow might need three judgments:
- Which team should handle this?
- How urgent is it?
- Does it require human attention?
A generative model can answer those questions. But if the application only needs a category, an urgency score, and a flag, generating language may be unnecessary overhead.
Jev focuses on those judgments directly.
TypeSafe calls this a “System One” model, borrowing the language of fast, intuitive thinking. The analogy is useful, but it is not a claim that the model thinks like a human.
The important distinction is architectural: bounded decisions rather than open-ended generation.
That helps explain the speed. It does not mean every Jev request finishes in under 500 milliseconds, or that Jev beats every specialized classifier. Actual performance depends on the input, deployment, network, and surrounding application.
What questions can you ask?
The best Jev questions have a clear purpose and a bounded answer space.
You supply the information to evaluate, the question, and the allowed outcomes. Jev returns probabilities or scores; your application determines what happens next.
Three useful question patterns are:
| Question type | Example | What you receive |
|---|---|---|
| Yes/no proposition | “Does this email request a refund?” | A probability that the proposition is true |
| Choice among options | “Which team should handle this: billing, technical support, or sales?” | Probabilities across the supplied options |
| Ordered score | “How urgent is this: low, medium, high, or critical?” | A score or probabilities associated with the defined levels |
You can combine these patterns. One customer message might be evaluated for topic, urgency, cancellation risk, and the need for human review.
Questions that fit
Here are illustrative questions you could build into a workflow:
| Use case | Question to ask | Possible outcomes |
|---|---|---|
| Email management | “What kind of message is this?” | Client request, sales inquiry, newsletter, internal update, other |
| Support triage | “What is the appropriate handling path?” | Standard workflow, specialist queue, human escalation |
| Customer retention | “Does this message suggest an intention to cancel?” | Probability of cancellation intent |
| Lead qualification | “How well does this inquiry match our stated service criteria?” | Strong fit, possible fit, poor fit |
| Document screening | “Is this document relevant to the research topic provided?” | Relevance score |
| Content planning | “Does this proposed article substantially overlap with our existing content?” | Probability of overlap |
| Video selection | “How strong is this transcript segment as a standalone clip?” | Ordered quality score |
| Agent oversight | “What review level does this proposed action require?” | Routine, confirmation required, block |
| Model routing | “Which available model tier fits this request?” | Small model, larger model, specialist workflow |
| Browser interaction | “Which listed page element matches the user’s instruction?” | One of the supplied elements, or no match |
These are not invitations for Jev to invent an answer. They are instructions to judge the information against choices or criteria you define.
For example, “qualified lead” needs a definition. Otherwise, the model is interpreting a vague label rather than applying your business requirements.
Questions that do not fit
Some requests require something fundamentally different:
- “Write a customer response.”
- “Explain why this strategy failed.”
- “Create a new marketing campaign.”
- “Summarize this report.”
- “Calculate the exact financial exposure.”
Writing and explanation belong with a generative model. Exact calculations belong with deterministic tools.
A useful test is:
Do I need new content—or a judgment among known possibilities?
If you need a judgment, Jev may fit. If you need an explanation, calculation, or new piece of writing, use the appropriate tool instead.
Useful without generated text
Jev can be useful even when there is no generative model anywhere in the final workflow.
The output might be a label in a spreadsheet, a priority flag in a ticketing system, a filtered document list, or a decision to request approval.
An inbox that prioritizes
Instead of writing summaries of every email, the system evaluates:
- Does this require a response?
- Does it contain a deadline?
- Is it a customer issue?
- Does it belong in today’s priority queue?
The visible result is an organized inbox—not AI-generated prose.
A research queue that filters
A researcher supplies a topic and candidate document excerpts.
Jev scores relevance against the stated criteria. The application presents the most relevant candidates for review.
The model is not performing the research or proving the documents are correct. It is helping allocate attention.
A content pipeline that screens
A writer supplies proposed article ideas alongside existing titles and descriptions.
The workflow asks:
- Is this idea meaningfully different?
- Does it fit our intended audience?
- Is it primarily technical, strategic, or introductory?
- Should it enter the editorial review queue?
The result is a shortlist. A person or generative model can handle drafting later.
An agent that pauses
An agent proposes an action. Jev evaluates whether it appears routine, sensitive, or risky.
The application then decides whether to proceed, request confirmation, or block the action.
That classification is an additional safeguard—not a replacement for permissions, security controls, or human approval.
None of these examples requires a chatbot response. The intelligence becomes part of the application’s behavior.
And “no generated text” is not the same as “no-code.” Jev still needs an application, automation platform, or integration to turn its judgments into useful actions. The supplied demonstrations include agent-assisted setup and workflow integrations, rather than just a standalone chat interface.
Micro-trading in under 500ms?
One of the most striking demonstrations comes from Lewis Jackson Investing’s video, I Gave JEV Control of a Trading Bot.
The demonstration connects a Jev-based decision loop to Bitcoin paper trading through Alpaca. The surrounding system supplies market information and strategy conditions, while the decision layer evaluates bounded outcomes.
Near 14:53, the presenter reports approximately 395 milliseconds, describing it as time “between trades, between decisions.”
That is an impressive illustration of a sub-500-millisecond decision loop. But it needs careful wording:
The video reports approximately 395 milliseconds between decisions in a paper-trading demonstration. It does not establish that live trades were consistently executed and filled within that time.
Earlier, the presenter describes account activity occurring roughly every three to five seconds. The video does not provide an independently verified breakdown separating inference time, order submission, broker acknowledgment, and execution.
Those distinctions matter.
A fast trading system must do more than make a prediction. It must ingest fresh data, evaluate the strategy, apply risk limits, submit an order, and handle execution. A fast decision is only one component.
The 300ms engineering target
Alex Hitt’s Jev Trader GitHub on Monad: Under a 300ms discusses a related engineering challenge: fitting an AI-assisted trading pipeline into a very tight processing budget.
The video examines proposed improvements such as local feature extraction, inventory management, transaction-signing choices, and circuit breakers that reject stale decisions.
Its useful lesson is broader than trading:
If the information has changed before the decision arrives, a fast answer can still be too late.
The 300-millisecond figure should be treated as the video’s optimization target, not an independently demonstrated guarantee.
Fast does not mean profitable
Sub-second decision-making does not establish a profitable strategy.
The Lewis Jackson demonstration uses paper trading and describes its initial strategy as generic. Andrew Warner’s 9 things you’ll actually do with Jev also includes a Bitcoin-signal example that performs poorly and cautions against blindly trusting the model with a portfolio.
A model’s probability for a trading-related proposition is not automatically the probability that a trade will make money.
Fees, spreads, slippage, stale data, inventory exposure, and execution failures can overwhelm any apparent advantage.
The remarkable part is the possibility of placing probabilistic judgment inside a rapid automated loop—not proof of a profitable trading system.
That is also the larger enterprise opportunity. Many applications need to interpret messy information and choose the next step, without writing anything at all.
The question is not whether every workflow needs more generated text.
It is whether some workflows would benefit from faster, better-defined decisions.
Alternatives: Jev is not alone
Jev is not the only way to make AI decisions without generating text. Alternatives such as GLiClass classify information against supplied labels, while SetFit builds efficient, specialized classifiers from relatively small sets of labeled examples. Conventional LLMs can also return structured decisions, although they still generate tokens to produce those results. These approaches overlap with Jev’s use cases, but they differ in flexibility, training requirements, and deployment options.
| Approach | How it works | Example use | Generates text? | Key distinction |
|---|---|---|---|---|
| Jev | Evaluates context against typed questions and returns probabilities or scores | Assess urgency, route a request, or flag an action for review | No free-form text | A decision-focused model and interface for software workflows. |
| GLiClass | Classifies text against labels supplied at inference time | Categorize messages as billing, support, sales, or other | No | Supports zero-shot classification without task-specific training. |
| SetFit | Trains a classifier using sentence representations and labeled examples | Route tickets into established business queues | No | Suits recurring tasks where examples are available. |
| NLI-based classifiers | Evaluate how well text supports candidate label descriptions | Detect whether a message expresses a complaint or request | No | An established zero-shot classification approach. |
| LLMs with structured outputs | Generate answers constrained to a schema or allowed values | Return a category and urgency level in JSON | Yes, even without prose | Structured output constrains generation; it does not eliminate it. |
These are alternatives for overlapping tasks, not necessarily drop-in replacements. A classification score should not automatically be treated as a calibrated probability, and speed comparisons need the same inputs, hardware conditions, and measurement boundaries.
Is Jev’s approach entirely new?
No. AI systems have classified information, scored inputs, and selected among predefined outcomes for decades. Jev did not invent the idea of making decisions without writing prose. Its underlying purpose belongs to a long history of discriminative models—systems designed to distinguish among possibilities rather than generate open-ended content.
What distinguishes Jev is the way TypeSafe brings those capabilities together: a typed-question interface, parallel evaluation, probability outputs, and training focused on calibrated decisions. TypeSafe positions that combination as a “System One” model built for automation. These are meaningful design choices, but they should be distinguished from the broader, established field of classification.
The useful question for an enterprise is therefore not simply, “Is this a new category of AI?” It is, “Does this approach make our particular decisions faster, more reliable, or easier to integrate?”
Jev’s significance is not that AI can finally decide without writing. It is that decision-making—not conversation—is the product’s central design goal.
I have seen first hand and I’ve written many times before, the goal is not to use the newest model for everything. It is to choose the model that fits the job. Bottom line: use the right model for the task at hand.
Resources and source notes
- TypeSafe AI — Introducing System One Models & Jev. The company’s explanation of decision-native outputs, parallel evaluation, calibrated-decision training, and workload-specific performance claims.[typesafe]
- LangChain — Building a Harness with Jev. Explains typed questions, probabilities, classification, and Jev’s role in agent workflows.[langchain]
- IBM Technology — What Is Jev? The AI Model That Doesn’t Generate Text, October 1, 2026. Martin Keen explains question types, calibration, support routing, limitations, and complementary use with LLMs.[youtube]
- Lewis Jackson Investing — I Gave JEV Control of a Trading Bot, September 21, 2026. Bitcoin paper-trading demonstration; the approximately 395-millisecond claim appears near 14:53 and is described as time between trades or decisions, not independently verified live execution latency.[youtube]
- Alex Hitt — Jev Trader GitHub on Monad: Under a 300ms, September 19, 2026. Discusses proposed low-latency architecture improvements and a 300-millisecond processing target, including stale-decision controls.[youtube]
- Nate B. Jones — Why Developers Are Losing Their Minds Over AI That Can’t Write, September 21, 2026. Explores classification, workflow orchestration, spreadsheet judgments, and agent-assisted setup.[youtube]
- The Next New Thing — 9 Things You’ll Actually Do With Jev, September 26, 2026. Andrew Warner reviews practical demonstrations, including email classification, lead qualification, model routing, transcript scoring, automation integrations, and a trading example that performs poorly.[youtube]







