Let me as you a question, what if your AI application didn’t need another chatbot? 

Imagine an agent receiving a support ticket. It doesn’t need an essay. It needs to answer:

  • Which team should handle this?

  • Should this ticket be escalated?

  • Is the customer asking for a refund?

  • Which tool should run next?

  • Is the generated answer safe to send?

Traditionally, developers reach for an LLM, ask it to return JSON, parse the response, handle malformed output, and hope the model’s confidence means something.

Jev takes a different approach.

TypeSafe AI describes Jev as a System One model give it a state and typed questions, and it returns structured decisions such as a choice, score, or yes/no probability.

But there’s a catch, Jev is closed-weight and hosted.

That has created an interesting open-source ecosystem around the same idea. I went through eight projects that explore different ways to reproduce, approximate, or specialize in the Jev-style decision model.

First: What Exactly Is a Jev-Style Model?

A normal LLM works roughly like this:

Prompt
   ↓
LLM
   ↓
Generated text
   ↓
JSON parser
   ↓
Application logic

A Jev-style system changes the architecture:

State + Questions
       ↓
 Decision Model
       ↓
Typed probabilities
       ↓
Application logic

Instead of giving the statement like, “Read this ticket and explain which department should handle it.”

you ask:

department:
  type: choice
  options:
    - billing
    - technical
    - sales
    - support

The model only needs to select among those options.

That makes these models particularly interesting for routing, classification, scoring, guardrails, ranking, tool selection and agent control loops.

Now let’s look at the open-source alternatives.

1. SemIf

SemIf is one of the most interesting approaches because it doesn’t try to reproduce Jev’s undisclosed architecture.

SemIf uses open models and extracts typed option probabilities directly rather than generating an answer sentence and parsing it afterward.

The project targets consumer hardware, including an RTX 3090-class GPU, and also provides browser/WebGPU experimentation.

Why it matters

The architecture is surprisingly simple:

State
 ↓
Open Model
 ↓
Option Scores
 ↓
Probability Distribution
 ↓
Decision
  • No generated JSON.

  • No retry loop.

  • No regex parsing.

For developers experimenting with local AI agents, this makes SemIf particularly interesting as a practical starting point.

2. Bespoke Nimble 9B

Bespoke Nimble is arguably one of the most educational projects in this space. Instead of only publishing weights, Bespoke Labs published:

  • the model

  • training data methodology

  • training recipe

  • serving code

  • evaluation methodology

Nimble is a LoRA fine-tune of Qwen3.5–9B designed for typed decisions. It supports choice and boolean-style decisions and returns probabilities for the allowed answers.

The interesting part is the contrastive data curation technique.

The training examples deliberately modify facts so that changing a small piece of information can flip the correct decision.

The authors report 90.12% on their 324-example holdout, compared with 93.21% for Jev 1.13.0 on that particular evaluation. They also explicitly note that the evaluation is narrow and synthetic.

That caveat matters. A benchmark result is not the same thing as general intelligence. But the project gives developers something even more valuable a recipe for building their own decision model.

3. Decider

If you want a more complete System One research project, Decider deserves attention.

Its philosophy is straightforward “Don’t generate text. Make a decision.”

Decider supports typed choices, scores and yes/no-style decisions, with models ranging from 0.8B to 35B parameters.

The 2B model is built on Qwen3.5–2B, while the larger model uses Qwen3.5–35B-A3B.

It also includes:

  • calibration metrics

  • Brier scores

  • ECE

  • multiple datasets

  • browser tasks

  • games

  • Super Mario experiments

  • CPU support

  • CUDA support

  • Apple Silicon/MPS support

  • an HTTP server

That’s important because the project isn’t simply:

Model → prediction

It explores the entire decision-engine stack.

For engineers interested in calibration, inference performance and production architecture, Decider is particularly useful to study.

4. openjev

This project takes another route. Instead of turning an LLM into a generic choice engine, openjev fine-tunes Qwen3.5 into an NLI-style cross-encoder.

Its fundamental operation is:

Premise + Hypothesis
        ↓
Entailment / Contradiction / Neutral

The model card describes applications including reranking, grading, content guarding and game decision-making.

The newer 4B version also explores vision-based decision-making and game environments such as Doom and Minecraft.

This is an important architectural idea.

Sometimes your application doesn’t need “Choose option A, B or C”, it needs “Does this candidate logically follow from the current state?”

That’s essentially an NLI problem.

And that primitive can become a surprisingly powerful building block.

5. NanoJev

The name gives away the idea. NanoJev is a 0.6B decision model.

It uses a Qwen3–0.6B backbone with dedicated decision heads and produces probability distributions over supplied candidates without output-token decoding.

The project is especially interesting because it publishes the whole pipeline:

Dataset
   ↓
Training
   ↓
Checkpoint
   ↓
Inference
   ↓
Evaluation
   ↓
Replayable experiments

Its current unified model covers:

  • Maze & Snake

  • ViZDoom Basic

  • ViZDoom Predict Position

The authors report 128/128 successes for the Basic ViZDoom test and 27/128 for Predict Position on their matched test setup.

Those numbers should be treated as project-specific experimental results, not proof that NanoJev is universally better than Jev.

But the engineering lesson is more interesting A decision model doesn’t necessarily need billions of parameters.

For narrow, well-defined control loops, a small specialized model can be extremely compelling.

💡 Enjoying this article?
Every week day, I publish practical, production-ready deep dives covering Web development, System Design, Open source projects, Tech industry trends and AI Engineering and tools.

6. Laya

Laya takes the idea in a different direction multilingual, non-autoregressive decision making.

It provides multiple checkpoints, including:

  • ModernBERT-large - 421M

  • mmBERT-base - 322M

  • typed-decision variants

and supports more than 100 languages according to the project documentation.

The API exposes typed questions such as:

questions = {
    "department": {
        "type": "choice",
        "instructions": "Which department should handle this request?"
    }
}

That makes Laya interesting for applications such as:

  • multilingual ticket routing

  • email classification

  • customer-support automation

  • content filtering

  • agent routing

The project also focuses heavily on low latency and batch inference.

If your application needs local, multilingual decision infrastructure, Laya is a particularly relevant project to investigate.

7. OpenSourceJev

OpenSourceJev goes even closer to the hardware. The project describes itself as a local System One research experiment powered by llama.cpp and Qwen.

Its goal is to transform open-weight models into deterministic, strictly typed decision engines.

The architecture is roughly:

Prompt
  ↓
llama.cpp
  ↓
Model logits
  ↓
Candidate probabilities
  ↓
Typed decision

The interesting part is that the system works at the logit level rather than relying on generated text.

That makes it useful for developers interested in:

  • llama.cpp

  • local inference

  • logits

  • probability extraction

  • quantized models

  • CPU/GPU experimentation

It is more of a research implementation than a polished enterprise platform, but that’s exactly why it is worth reading.

8. Sales Conversion Model: A Specialized Alternative

This one needs an important clarification. It isn’t a general-purpose Jev replacement.

It is a specialized decision model for sales conversations.

The project uses reinforcement learning and PPO to predict sales-conversion outcomes from conversation data. Its model card describes synthetic training data containing more than 100,000 sales conversations and reports an 85ms inference time in its own evaluation.

This illustrates an important alternative to building a general Jev clone:

❝

Don’t build a general decision model if your problem is already highly specialized.

For example:

Sales conversation
       ↓
Specialized model
       ↓
Conversion probability
       ↓
Sales action

For a sales platform, that may be more useful than a general-purpose decision engine.

How These 8 Projects Differ

The projects also differ considerably in what their benchmarks actually demonstrate. For example, NanoJev’s headline numbers come from specific game environments, while Bespoke Nimble’s reported holdout is a narrow synthetic evaluation. Laya publishes its own performance measurements and limitations, while SemIf emphasizes reproducibility and direct scoring rather than claiming to reproduce Jev’s undisclosed training.

So don’t treat a single leaderboard number as the answer.

The Bigger Idea: AI Doesn’t Always Need to Generate Text

This is the part I find most interesting. For years, we treated every AI problem as:

Prompt → LLM → Text

Then developers started wrapping that text in JSON.

Prompt
 ↓
LLM
 ↓
JSON
 ↓
Parser
 ↓
Validation
 ↓
Retry
 ↓
Application

System One-style models suggest another architecture:

State
 ↓
Typed Question
 ↓
Decision Model
 ↓
Probability
 ↓
Code

That’s a fundamentally different way of thinking about AI engineering.

Your LLM can remain responsible for:

  • reasoning

  • writing

  • summarization

  • planning

  • coding

while a small decision model handles:

  • routing

  • filtering

  • ranking

  • tool selection

  • risk gates

  • classification

  • escalation

The future AI stack may not be one giant model.

It may be a collection of specialized models, each responsible for a very specific cognitive primitive.

Which One Should You Experiment With?

Don’t choose based purely on benchmark numbers. Choose based on your workload.

Want to understand the fundamentals?

Start with SemIf or NanoJev.

Want to learn how to train one?

Study Bespoke Nimble.

Want a serious research platform?

Look at Decider.

Need multilingual decisions?

Look at Laya.

Want to experiment with logits and llama.cpp?

Try OpenSourceJev.

Need semantic entailment/reranking?

Explore openjev.

Building a specialized sales system?

The Sales Conversion Model is the more focused approach.

Final Takeaway

Jev introduced an interesting idea AI doesn’t always need to talk. Sometimes it just needs to decide.

The open-source community is now exploring that idea from several directions:

  • frozen-model scoring

  • fine-tuned decision models

  • NLI cross-encoders

  • tiny 0.6B models

  • multilingual encoders

  • reinforcement learning

  • direct logit inference

  • complete open training pipelines

None of these should automatically be described as the open-source Jev. Jev’s internal architecture and training remain proprietary. What these projects provide are open implementations of the broader System One / typed-decision pattern, each with different trade-offs.

And that’s arguably more exciting than a single clone. Because now developers can experiment with the architecture themselves. The interesting question isn’t “How do I replace my LLM?”

It’s:

❝

“Which decisions in my application never needed an LLM in the first place?”

That’s where Jev-style models become genuinely interesting.

Benchmarks for this all models: https://benchmarkheaven.com/jev-models

Thank You for Reading!

I hope you found it helpful and informative. If you have any questions or feedback, feel free to leave a comment below. Your support and engagement mean a lot to me.

Happy Coding!

Reply

Avatar

or to participate