Open alternatives to Jev
The open System One models you can run yourself
Jev is a week old and has been hard to avoid since launch. TypeSafe's "System One" model is built to be called by software. You post one state plus a map of typed questions, and you get back typed values and calibrated probability distributions in a single forward pass. There are no generated tokens, nothing to parse, and no retry when the model decides to explain itself first. It caught on because it solves a problem most of us have been working around with constrained decoding, JSON schemas and regexes, because it costs $0.042 per million input tokens with output free, and because TypeSafe's docs are more candid about the model's failure modes than most vendors manage.
The interface is also easy to reimplement, and a lot of people have. This post catalogues the thirteen projects that clear the bar set out under "How I picked what to include". Roughly forty further candidates turned up in a sweep and did not clear it. Eleven of the thirteen run on an M1 Max with 64 GB of unified memory, which is the machine a follow-up post will use; NanoJev and djev-dev are documented rather than run, because leaving them out would misrepresent the field.
I summarise what each project does and where it matches or differs from Jev. I do not rank them: every number here was measured by someone else, and I say who each time. A second post will measure a subset on one machine with one suite, which is the only way I know to turn a catalogue into a comparison.
Interestingly, two of the most methodologically careful projects here, AnyJev and open-alternative-jev, train nothing at all. Both wrap a stock open model, read its option logits properly, and reach Jev's accuracy with better-calibrated probabilities. Both are single-benchmark, single-model-family results, so this settles nothing, but it does change what you should be asking of the nine trained open decision models.
→ Jump straight to the results
What is Jev?
This section covers only the details the comparisons depend on. TypeSafe's own docs and the four independent analyses in the appendix cover the rest.
There are three primitives (V):
| Primitive | You supply | You get back |
|---|---|---|
noul |
optional descriptions of what true and false mean | noul: P(yes) as a float. No confidence field. |
choice |
option → rubric description, max 255 options | choice (argmax), probabilities, confidence |
score |
an ordered array of 2–10 level descriptions | score (probability-weighted, lands between levels), legend, probabilities, confidence |
Post-training is described as RLCD, reinforcement learning for calibrated decisions: an outcome-based objective against proper scoring rules rather than RLHF against human preference. The commercial envelope (V): model jev-1.13.0, $0.042 per million input tokens with output free, 64k context of which 32k is state plus the single longest question, text only, English first, and no fine-tuning or customisation of any kind. That last row matters more than it looks, because every open competitor's main lever is fine-tuning on your own labels.
confidence is a peakedness statistic, not a probability of being correct. For a choice over K options it is (K · p_max − 1) / (K − 1), normalised so a uniform distribution gives 0 and a one-hot gives 1. It says nothing about whether the argmax is right. TypeSafe say this themselves and return the full probabilities vector so you can compute your own, and two independent measurements suggest you should: anth.us finds AUROC 0.83 for ranking correct from incorrect but 7–15 points of overconfidence as a probability, and scienthoon finds the field "was never better than the max probability and sometimes much worse" (both T).
The vendor's own jaggedness page lists nine failure modes: literal reading of implied intent, arithmetic and counting, date comparison, indirection, large irrelevant state acting as a distractor, injected content, contradictory instructions, no guaranteed invariants between related questions, and generation. Their guidance — move arithmetic into code, filter the state, compare dates in code — amounts to saying Jev is a semantic verifier rather than a reasoner. I think that is a fair description and a reasonable thing to sell.
Why you would look for alternatives to Jev?
Cost is not the reason. At $0.042/Mtok input with output free, and JevBench's measured ~950 input tokens per decision (V + T), that is roughly 4 cents per 1,000 decisions. scienthoon's entire 4,621-call study cost about 6 cents. Check the arithmetic before you self-host to save money.
Four reasons that do hold up:
Trainability. Jev offers no fine-tuning, no LoRA and no per-account weights (V). You adapt through
stateandcriteriaonly. Nine of the thirteen projects below take your labels, and nimble's jump from 66% to 90% on 2,676 examples shows how cheaply that can pay.Latency. Hosted Jev is around 250 ms wall clock from a US or European origin, consistent across three independent measurements, against roughly 160 ms of upstream service time. A local encoder is 13–130 ms with no network involved.
Data residency. Zero data retention is an enterprise option. Running locally keeps the data on your own hardware.
Pinning and continuity. Aliases move, the docs say rate limits can change without notice (V), and the Vercel AI Gateway route cannot pin a version at all. That matters once you have tuned thresholds against one build.
The counterweight is that Jev is still ahead on knowledge-heavy decisions, which is where decider and kev both report their largest gaps, and that the alternatives were all days old when I wrote this.
How I picked what to include
A sweep of GitHub topic:system-one and topic:jev, a full-text repository search and a third-party index turned up roughly 40 projects beyond the obvious ones, plus about 20 "awesome-jev" link collections. Four criteria, all of which must hold:
It implements the System One contract: typed primitives with a probability distribution per question, read from the model rather than generated as text. A service that calls the hosted Jev API is an application, not a comparator.
It runs locally and reproducibly, with open weights or a stock open-weight base plus published inference code.
It is pinned and licensed.
It has been measured, or ships a bench that can measure it.
One warning about discovery aids: the widely-linked index at systemonemodels.org lists kev, decider and SemIf as having no published metrics, which is wrong for all three. Use it to find things, not to cite them.
Evidence grading
I grade every number and don't mix classes inside a comparison:
| Class | Meaning |
|---|---|
| V | Verified spec — read from vendor docs, repo code or manifests |
| T | Third-party measurement — someone independent ran it and published raw data |
| S | Self-reported — the authors measured their own model against a number they did not produce |
Almost every "X vs Jev" table in these repos compares a class S number against a class T or S number, measured on different items, prompts, dates and hardware. Several of the repos say so themselves. Pooling those numbers into one ranking would produce a ranking that is confidently wrong.
What was excluded, and why
| Excluded | Reason |
|---|---|
| openjev-sglang | No licence file (criterion 3). Also 35B |
| openJev-verdict-2.0 | Verdict 2.0's weights are unobtainable (criterion 2) |
| jevlike | No accuracy or calibration metric (criterion 4). A teaching repo, and a good one |
| PlayJev, dohnuts, jev-visual | Multimodal or GUI-game specialists (criterion 1 in spirit) |
| notjev, SiliconLabAI/OpenJev, iamaamir/system-one | Provider-neutral wrappers over someone else's endpoint (criterion 1) |
| ~13 small repos: mini-jev, litjev, jevmlx, fastjev, minojev, jevfire and others | Created 16–22 September, no published metric and no runnable bench (criterion 4). Several look plausible; none is yet measurable |
The alternative field
The first table says what a project is; the second says whether you can use it. Row order is the same in both.
What each alternative is
| Project | Created | Architecture | Params | Licence |
|---|---|---|---|---|
| Jev 1.13.0 (reference) | — | inferred sparse MoE decoder + linear head | ~10B active (inferred, T) | commercial, hosted only |
| laya | 18 Sep | ModernBERT-large / mmBERT-base encoders, three checkpoints + router | 421M / 322M / 421M | Apache-2.0 |
| laya-mlx | 19 Sep | MLX port of the three laya checkpoints | same | Apache-2.0 |
| kev | 17 Sep | Qwen3.5-Base + LoRA + pointer head | 0.8B / 4B / 9B | Apache-2.0 |
| SemIf | 16 Sep | stock Qwen3.5-4B read through option logits, no training | 4B | MIT (code) |
| GLiNER2 | Jul 2025 | encoder, schema-driven extraction and classification | ~300M | Apache-2.0 |
| NanoJev | 17 Sep | Qwen3-0.6B + decision heads | 0.6B | MIT (repo); HF repo unset |
| nimble | 18 Sep | Qwen3.5-9B + LoRA r16 | 9B | card says Apache-2.0; repo has no LICENSE file |
| von | 18 Sep | ModernBERT-Large bidirectional encoder | 395M | Apache-2.0 |
| decider | 16 Sep | Qwen3.5-Base + decision head | 1.9B / 34.7B-A3B (3B active) | Apache-2.0 |
| jeff | 19 Sep | GLiFormer-large-v1 encoder | 400M | MIT |
| AnyJev | 21 Sep | readout layer over any LLM, no training | any | Apache-2.0 |
| open-alternative-jev | 18 Sep | packed-state readout over any open LLM, no training | any | Apache-2.0 |
| djev-dev | Sep | DiffusionGemma-26B-A4B via vLLM | 26B MoE, 4B active | Apache-2.0 |
Star counts are a poor guide as of 23 Sep 2026. laya gained 2,583 stars in the 24 hours between my two research passes, and the two most methodologically careful projects in the list have 68 and 51 stars. Sort by something else.
Whether you can use it
These columns decide whether you can adopt a project. They are currently scattered across thirteen READMEs, and collecting them is useful.
| Project | Primitives | Max choice options | Context | Drop-in /v1/systemone? |
Apple Silicon |
|---|---|---|---|---|---|
| Jev 1.13.0 (reference) | noul, choice, score |
255 | 64k (32k state + longest question) | it is the spec | hosted only |
| laya | all three | no hard cap, degrades badly above ~20 | 512 English, 1,024 multilingual (state shares it with the options) | no — own Python API | CPU/MPS, 193–464 ms CPU |
| laya-mlx | all three | as laya | as laya, per checkpoint | no — own Python API | native MLX, 13.4 ms P50 (M3 Max) |
| kev | all three | 1–255 | 8,192 serving (state + one question); trained at ≤384 state tokens | yes, native | yes, but slow on the Qwen3.5 generation |
| SemIf | choice only — a yes/no is two options | not published | base model's (Qwen3.5-4B) | no — own scorer | yes: MLX, MPS and llama.cpp |
| GLiNER2 | n/a — extraction schema | not published | encoder window; long documents go through a 384-word chunker | no | CPU-friendly |
| NanoJev | all three | 2–255 candidates | not published | no — own /api/evaluate |
no — CUDA + ViZDoom |
| nimble | all three | ≤26 (A–Z encoding) | 2,048-token prompt cap | no — own ParallelScorer |
yes, MLX first-class |
| von | all three | not published — two_stage_choice exists for taxonomies above ~25 |
not published | yes, native | yes, MPS |
| decider | all three | 2–255 | 32k | yes, native | yes, MPS merged 22 Sep |
| jeff | all three | 64 (JEFF_MAX_LABELS, configurable) |
20,000 state characters (JEFF_MAX_STATE_CHARS); no token figure published |
yes, native | yes — CUDA → MPS → CPU |
| AnyJev | all three via readout | 26 | base model's | not yet — on the roadmap | via transformers/MPS |
| open-alternative-jev | all three — Choice, yes_no, scale |
26 (A–Z, each letter one token) | base model's | no — own packed API | via transformers |
| djev-dev | all three, plus images as options | not published | not published | no — Jev-shaped body at its own /v1/request |
no — CUDA-only vLLM |
"Not published" means I went through the repo and the figure genuinely isn't there. Six of the thirteen don't state an option cap and five don't state a context limit, and those are exactly the numbers that decide whether you can adopt something. Most of these repos publish an accuracy number against Jev before they publish their limits.
Two things stand out:
kev, von, decider and jeff implement TypeSafe's actual
/v1/systemonewire format, so the official SDK points at them with an environment variable change. For adoption that matters more than any accuracy number. kev's CI runs TypeSafe's own example requests against its server.Check the option cap before you read the accuracy table. A 40-option routing problem rules out nimble, AnyJev and open-alternative-jev immediately, degrades laya, and is over jeff's default limit, whatever those projects score.
The projects
Each project follows the same template, and carries the repo link plus the exact commit and checkpoint revisions I read, because in a field this young an unpinned claim stops being checkable within days. Those pins are what I read, not what I recommend you run.
1. laya
Read 23 Sep 2026 at 725e020 · weights convaiinnovations/laya@1c5edc1, convaiinnovations/laya-typed-decisions@f9ab0b2
Three encoder checkpoints — ModernBERT-large 421M, mmBERT-base 322M, and a typed-decisions fine-tune — with a script-detecting router that picks between them in under 0.5 ms. 2.4 GB total, Apache-2.0.
laya documents its own limits more openly than most of the field does. The base checkpoints score near chance zero-shot on its own typed-decisions benchmark: 0.362 and 0.342, against a 0.318 random baseline and a 0.461 majority-class floor. The headline 0.766 comes from the checkpoint fine-tuned on that benchmark's own training split, and the authors say so (S).
Three other limits. Raw ECE is 0.466 and falls to 0.081 only after per-(type, option-count) temperature fitting, and laya-multilingual ships with no fitted temperatures. High-cardinality choice collapses — Banking77 0.425 against Jev's 0.870 — because ~77 labels share a 192–256 token head budget, about three to four tokens per label. Khmer scores 0.000 accuracy at 0.952 confidence, which is why the router exists.
Differs from Jev: encoder rather than decoder, multilingual, and designed to be specialised rather than used zero-shot. The authors call it "a fast base to specialise, not a zero-shot decision engine", which is the fairest self-description in this field.
Runs on: Mac (CPU/MPS) at 193–464 ms on CPU, 32.8 ms on a T4. It ships a Kaggle 2×T4 fine-tuning notebook.
2. laya-mlx
Read 23 Sep 2026 at 0a85951 · weights aac6fef/laya-mlx@20aed81
An independent MLX port of all three laya checkpoints — encoder, decision transformer, scoring head and action head — with no PyTorch and no transformers runtime.
On an M3 Max in FP16, end to end including tokenisation: 13.42 ms P50 for one short question, 7.39 ms for the multilingual checkpoint, 146.8 and 395.0 questions per second at batch 50, under 1 GiB peak allocation (S, but unusually well documented).
Port fidelity is verified rather than asserted. All three checkpoints matched the upstream selected answer on 63/63 validation questions in both FP32 and FP16 — 378/378 comparisons — with pinned revisions, recorded weight hashes and 36 checksum-verified files. Almost no other port in this field documents fidelity at all.
Differs from Jev: it inherits laya's architecture and laya's weakness. It sets the latency floor for local decisions, but those decisions come from checkpoints that score near chance zero-shot, so latency is not the constraint that matters here.
Runs on: Apple Silicon only, macOS 14+.
3. kev
Read 23 Sep 2026 at 557598f · weights jaredpalmer/kev-4b@485ace8, jaredpalmer/kev-0.8b@54f4f87
Qwen3.5-Base plus LoRA plus a pointer head that scores each option's </opt> hidden state against the question's <decide> hidden state. Three sizes: 0.8B, 4B and 9B.
Its harness, more than its model, makes it the most evaluation-ready project here: a genuine TypeSafe-compatible /v1/systemone, a benchmark runner that scores any System One endpoint over --remote, a live-Jev runner, and paired bootstrap confidence intervals between two saved runs. Its evals/external/ already holds two other projects' test sets in this format with their published live-Jev results.
It also reports a metric I did not see elsewhere: share of decisions automatable at a 5% error budget, 0.45–0.57 against Jev's 0.70 (S, but paired and on frozen items). Confident errors are 4.0% against Jev's 3.7%, after a fitted temperature of 2.1–2.4 baked into the checkpoint. Knowledge is where it loses, and the base model sets that gap: MMLU 0.74 against 0.90.
Weakness that shapes the local shortlist: there are no fast Metal kernels for Qwen3.5's Gated DeltaNet layers, so PyTorch runs reference code. Median on an M5, five questions, ~230-token state: 329 ms for 0.8B, 779 ms for 4B, about 2 s for 9B. The previous Qwen3 generation is three to six times faster on the same Mac (4B at 174 ms). An MLX backend is the next planned change.
Runs on: Mac, slowly, and any CUDA box.
4. SemIf (formerly OpenJev)
github.com/TheoLeeCJ/SemIf-OpenJev
Read 23 Sep 2026 at 1f2dea3 · weights Qwen/Qwen3.5-4B@851bf6e
Stock Qwen3.5-4B with no training at all, read through option logits. MLX, PyTorch/MPS and llama.cpp backends, all added in the last week.
It ranks around #2 to #4 on JevBench depending on the revision you read (T). On one RTX 3090, same model and same state, 21 binary criteria took 1.023 s through direct typed logits against 5.332 s for the autoregressive JSON equivalent, a 5.2× difference (S). The authors note that the two readouts agreed on only 18 of 21 criteria, so this is a systems comparison rather than a claim of semantic equivalence.
Differs from Jev: nothing was trained. That makes SemIf the control condition this field needs. It replaces "which fine-tune is best" with the more useful question of whether any of this training beats reading a stock model properly.
Runs on: Mac natively, three backends, ~8 GB in bf16.
5. GLiNER2
Read 23 Sep 2026 at 4abb613 · weights fastino/gliner2.5-base-v1@1a8bc24
The one pre-Jev entry (July 2025): a ~300M encoder for structured extraction and classification. It was pulled into this conversation because it was already producing schema-driven typed outputs from an encoder.
Unlike every other project here, GLiNER2 has no published System One-framed comparison against Jev. I am listing it because it belongs in the lineage, not because there is a number to cite.
Differs from Jev: extraction-first rather than decision-first, and no calibrated probability contract over a typed question set.
Runs on: anything, including CPU.
6. NanoJev
github.com/TianyuCodings/NanoJev
Read 23 Sep 2026 at 76fdfc9 · weights C-Tianyu/NanoJev@047b927
Qwen3-0.6B with decision heads, trained on 18,760 decision questions across four games: ViZDoom Basic, ViZDoom Predict Position, a 50×50 maze, and Snake.
It beats Jev decisively on those games — ViZDoom Basic 128/128 against Jev's 56/128, maze solved in 225 attempts against 2,738 (S) — because it is a domain-specialist embodied-control model, not a general text-decision model. Benchmarking it against text decisions would be as misleading as testing Jev on Snake. Its sibling project JevHarness, which generates LLM-authored task-specific Jev harnesses, is arguably more relevant to using Jev well.
Runs on: CUDA only, 21.6 GB of checkpoints, plus ViZDoom.
7. nimble (Bespoke Labs)
github.com/bespokelabsai/nimble
Read 23 Sep 2026 at 38edc3b · weights bespokelabs/Bespoke-Nimble-9B@e93fabc
Qwen3.5-9B plus LoRA r16, built in a day, trained on 2,676 contrastive examples. The curation method is worth stealing: write two near-identical examples that differ in one fact that flips the label, verify both with separate model calls, then verify that removing either evidence sentence makes the fact unknowable.
Holdout results on 324 items: Jev 93.21%, Nimble-9B 90.12%, Qwen3.8-27B 84.88%, base Qwen3.5-9B 66.36% (S). That jump from 66% to 90% on 2,676 examples is the strongest single argument in the field that a small local model closes most of the gap cheaply.
The authors state their constraints plainly: a 2,048-token prompt cap, ≤26 enum options through A–Z encoding, and probabilities that are explicitly uncalibrated at temperature 1.0. The labels are synthetic and, in their words, no person has reviewed them.
There is also a licence discrepancy. The model card says Apache-2.0, but the repository has no LICENSE file. Resolve that before you depend on it.
Runs on: Mac with 64 GB — the README says so explicitly — at 444 ms median on an M5 Pro, 106 ms on an H100. ~18 GB base plus a 0.17 GB LoRA merged on CPU.
8. von
Read 23 Sep 2026 at 14be898 · weights wfzyx/von@d8bb5e0
ModernBERT-Large 395M bidirectional encoder, dual cross-entropy and Brier loss, 250k NLI examples plus ~290k multi-domain, a fitted temperature of 1.1692, the TypeSafe wire format, Python and TypeScript SDKs, and a claimed ~18 ms GPU latency at a 3.2 GB download.
The README's headline and its own table disagree. The bullet list claims "SOTA Empirical Accuracy: 91.23% ... surpassing published commercial alternatives". The benchmark table two paragraphs below shows von at 72.0% macro against Jev's 96.6% on the same 49-task suite, a 24-point deficit. The ViZDoom row (von 9.00 kills against Jev's 5.62) is real, but it is one narrow zero-shot control task. The document does not support its own framing.
von is architecturally interesting, cheap to run and plausibly among the fastest local options, but every published claim needs independent verification before anyone cites it.
Runs on: Mac (MPS), CUDA, ROCm, CPU. Python ≥3.12.
9. decider
Read 23 Sep 2026 at 7557fe0 · weights Mapika/decider-2b@fa996ce
Two released sizes on Qwen3.5-Base, 1.9B and 34.7B-A3B, plus a 0.8B and a vision variant. 32k context, 2–255 choice options, 2–10 score levels: it matches Jev's declared envelope more closely than anything else here (V).
It reports third-party rankings rather than its own (T). On JevBench: Jev #1 at 75.4, decider-35b-a3b #10 at 68.9, decider-2b #21 at 64.6. On Decision Index v0.1, decider-35b-a3b is #4 of 32 and the highest-scoring trained model; the two entries above it are zero-training wrappers on stock 27B-class models.
It names where it loses: knowledge 0.51 against Jev's 0.69 per-area, and decider-2b's top-label ECE on JevBench hard items is 0.30, meaning it is confident where it is wrong. Its "Limits, stated plainly" section is the most candid in this list. One line from it: "rules written into the question are not followed at this size" — a one-sentence question scores 0.67, a paragraph of rules 0.24.
Apple Silicon is measured rather than asserted, and this is the only project in the list that did so. On an M1 Pro in float16, median request is 133 ms with the MPS patch against 171 ms without, and the MPS path scores 0.7553 accuracy / 0.0438 ECE against the published bf16 row's 0.756 / 0.041 on 1,500 held-out examples. That is a numerical-parity check on exactly this hardware class.
Runs on: pip install decider-ai, scripts/serve.sh gives /v1/systemone. 1.9B is 3.8 GB; 34.7B is 69 GB in bf16 and will not fit in 64 GB unified memory.
10. jeff
github.com/logan-markewich/jeff
Read 23 Sep 2026 at 34b32f9 · weights knowledgator/gliformer-large-v1@d0a4e53
A self-hosted drop-in for the Jev API on GLiFormer-large-v1 (400M). Device order is CUDA → MPS → CPU, so it works on a Mac unmodified, and the official SDK points at it with one environment variable. All three primitives.
It publishes where it loses. Measured on 1,600 labelled items across eight datasets: AG News 75.5% against Jev's 90.5%. JevBench v1.2.2 score 66.9, #9 of 18, against Jev's 75.3 — and the author notes that this rank comes from cost, because jeff is #14 of 18 on intelligence (T). Latency is 151 ms p50 on an L4 via Modal against Jev's 129 ms.
A small encoder that is six times cheaper (~$2.6 against ~$15.6 per million single-question requests) and fifteen accuracy points worse is a real point on the trade-off curve, and no other project in this list occupies it.
Runs on: Mac unmodified, ~0.8 GB.
11. AnyJev (Nokia Applied Research)
github.com/nokia-applied-research/AnyJev
Read 23 Sep 2026 at 54f1b35 · no weights of its own
AnyJev is a readout layer over any LLM rather than a model, and it trains nothing. Its thesis is that the problem is the readout, not the weights: max_tokens=1 plus logprobs already gives you a ranking, but one that moves when you reorder the options, and a confidence number you cannot threshold on. Both are properties of the readout, and both are fixable without a single label.
There are three levels, and every returned decision carries its level so downstream code can refuse to act on the wrong one. raw is a restricted softmax over label tokens. L0 adds cyclic-shift marginalisation — show K options in K rotations, combine in log space — plus prior correction, with no labels. L1 adds temperature scaling on 100–500 labels.
Qwen3-8B on BANKING77 20-way, n=300 (S):
| raw | L0 | L1 | |
|---|---|---|---|
| Answer flips when options reversed | 0.230 | 0.073 | 0.077 |
| Accuracy | 0.747 | 0.803 | 0.807 |
| ECE | 0.240 | 0.184 | 0.095 |
| Auto-decidable at ≤5% error | 7.7% | 46.3% | 52.0% |
All of that is on frozen weights, and an independent rerun came back bit-identical on every zero-label number.
The limitations section is exemplary, and it constrains how those numbers should be read: L0 is not a free win on every split, the batch prior needs at least 8 items of the same question, 26 options maximum, L1 does not survive distribution shift, coverage at a 5% risk budget is a high-variance point estimate at n=300, and everything so far is one model family.
Cost: a K-option choice costs K prefills (2 for noul, 1 for score), all sharing the state prefix and batchable. About 0.25 s per decision at batch 32 on an H100 with 20 permutations. Whether that multiplier is affordable on an M1 Max is an open question.
Runs on: whatever your base model runs on. No weights of its own.
12. open-alternative-jev
github.com/ikermoel/open-alternative-jev
Read 23 Sep 2026 at 4a85df1 · no weights of its own
Zero training. The state is written once and every question is answered from the next-token distribution at its own packed position.
The packing result is clean (S). On RACE-H, 250 passages × 4 questions: packed scoring gives 92.9% accuracy at 4.55 q/s using 186,898 tokens, against 92.6% at 1.66 q/s and 468,583 tokens for one question per forward pass. That is 2.5× fewer tokens and 2.3× the throughput at the same accuracy.
The headline result is on the 2,000-decision LocalLLaMA/typed-decisions suite: zero-shot Qwen3.6-27B at 73.7% accuracy and ECE 0.020, against Jev's 72.7% and 0.144. Both sit at the teacher's 73.5% self-agreement ceiling, so the accuracy column is saturated and calibration is the only place left where the models differ. The authors verified their scorer by reproducing laya's published numbers exactly before running anything else, which is what makes the rest credible.
The caution that goes with that number: JevBench ranked this project using the author's own option order (A. yes, B. no) and found the same model scored 21% instead of 72% with the options reversed. Read it alongside AnyJev, which exists to remove exactly that failure mode.
Runs on: your base model. ≤4B bases on a Mac; the 27B result was measured on an H200 MIG slice.
13. djev-dev
Read 23 Sep 2026 at 3ce907e · weights google/diffusiongemma-26B-A4B-it@f7f5b7f
Documented, not runnable on a Mac. DiffusionGemma-26B-A4B through vLLM: it denoises all answer positions together rather than scoring a next-token distribution, with native image input and images as options. It adds no weights of its own.
It is the only architecturally different entry in this list, at around #5 on JevBench with 73.0 (T). A 26B MoE plus CUDA-only vLLM puts it out of reach here. If the comparison ever adds a "does the readout architecture matter" arm, this is the entry for it.
Where these converge, and where they split from Jev
Same as Jev. Typed primitives with probabilities read from the model rather than generated as text. One forward pass. Probabilities exposed so you can compute your own confidence. Four of thirteen match the wire format exactly.
Different from Jev, on four axes:
Trainability. Jev offers none. Nine of these can be fine-tuned on your own labels.
Calibration as shipped. kev bakes in T ≈ 2.1–2.4, von ships T = 1.1692, laya ships fitted temperatures for one checkpoint and none for another, nimble ships T = 1.0 untuned, and Jev ships whatever RLCD produced. Comparing raw ECE across these compares post-processing choices, not models. I think this is the single biggest reason the existing comparisons do not mean what they appear to mean.
Architecture. Encoders (laya, von, jeff, GLiNER2), decoder plus LoRA plus head (kev, nimble, decider, NanoJev), pure readout over a stock model (SemIf, AnyJev, open-alternative-jev), and diffusion (djev-dev). Jev is inferred to be a sparse MoE decoder with a linear head.
Scope. Jev is general text decisions. NanoJev is embodied control. GLiNER2 is extraction. djev-dev is multimodal.
The shared weakness: option order
Archer Hume measured a −0.28 mean log-odds shift in Jev from adding irrelevant options. JevBench recorded 72% against 21% from reversing two options. AnyJev measured a 0.230 raw flip rate. Nobody is immune, including Jev. Any comparison that runs each choice item in one option order is not measuring what it claims to measure.
Zero-training readouts are already competitive
AnyJev and open-alternative-jev, on different bases and different suites, converge on the same finding: a calibrated zero-training readout of a stock open model matches or beats Jev on the probabilities while matching or slightly trailing on accuracy. AnyJev moves auto-decidable traffic 6.8× on frozen weights. open-alternative-jev puts a zero-shot 27B at Jev's accuracy with seven times better ECE. AnyJev's re-measurement of laya on laya's own benchmark has the fine-tuned checkpoint winning on argmax (0.768) while its ECE is six times worse than a calibrated stock model's (0.215 against 0.036).
AnyJev also previews a closed-form head: a shrunk-LDA or ridge matrix fitted on the same hidden state from 200 labels per question, no gradients, seconds on a CPU, reaching 0.771 accuracy and 0.120 ECE against fine-tuned laya's 0.768 and 0.215.
Both results are single-model-family and single-benchmark, and both carry the order-sensitivity caveat above, so I would not treat this as settled. It does raise the question I want the second post to answer: if a properly-read stock model is already at Jev's level, what are the nine trained open decision models for?
Shortlist A — runs on an M1 Max with 64 GB
Criteria: open weights or a stock base, runs on MPS or MLX without patching, pinned and licensed, and has a published number I can check. If you want one answer rather than nine, start with decider-2b.
| Model | Download | Published Mac latency | Pick it when |
|---|---|---|---|
| decider-2b | 3.8 GB | 133 ms median (M1 Pro) | You want the closest drop-in to Jev's envelope, with Mac numbers someone actually measured rather than asserted |
| laya-mlx | 2.4 GB | 13.4 ms P50 (M3 Max) | Latency dominates and you are willing to fine-tune. Near-chance zero-shot, so this is a base to specialise, not a drop-in |
| SemIf (Qwen3.5-4B) | ~8 GB | 1.02 s / 21 criteria (3090) | You already run a Qwen 4B and would rather read it properly than add a model |
| AnyJev on a base you have | 0 | K prefills per K-option choice | You want calibrated, order-stable decisions from existing weights and can afford the prefill multiplier. 26-option cap |
| open-alternative-jev on a base you have | 0 | packing cuts tokens 2.5× | Many questions against one long state — packing is the cheapest way to ask them |
| jeff (GLiFormer 400M) | ~0.8 GB | 151 ms p50 (L4) | Cost dominates and you can absorb ~15 accuracy points. No other project here occupies this corner |
| kev-0.8b (or kev-4b@qwen3) | ~1.6 GB | 329 ms / 174 ms (M5) | You want the harness as much as the model. Use the previous-generation base on a Mac |
| von-1.0 | 3.2 GB | ~18 ms claimed | You want a cheap encoder and are prepared to verify its claims yourself |
| nimble-9b | ~18 GB | 444 ms median (M5 Pro 64 GB) | You have 64 GB, ≤26 options and prompts under 2,048 tokens, and you intend to fine-tune |
If you have labels, laya fine-tuned and nimble are the two built to be specialised, and both publish what specialisation bought them. On a smaller Mac, everything except nimble-9b fits in 16 GB.
One practical warning about dependencies: decider-ai pins numpy<2 and transformers>=5, laya wants transformers>=4.48, and laya-mlx wants neither. One virtualenv per model is mandatory. nimble's README already enforces three for the same reason.
These recommendations rest entirely on other people's numbers, and the second post is where I find out whether I would still give the same answer afterwards.
Shortlist B — needs a real GPU
| Model | Why it will not fit | What it would answer |
|---|---|---|
| decider-35b-a3b | 69 GB bf16, exceeds 64 GB unified; the NVFP4 build is Blackwell-targeted | Does scale close the knowledge gap for a trained model? |
| open-alternative-jev on Qwen3.6-27B | 8-bit needs ~35 GB; measured on an H200 MIG slice | Does scale close it without training? |
| djev-dev | 26B MoE plus CUDA-only vLLM | Does the readout architecture itself matter? |
| kev-9b / nimble-9b for latency | They run on a Mac for correctness, but Mac latency numbers are not GPU latency numbers | Fair latency on known hardware |
| NanoJev | CUDA serve plus ViZDoom, and it is a games model | Nothing this comparison is asking |
None of these is required for the core comparison. One thing to know before you plan around free compute: two T4s do not pool into 32 GB, so Kaggle is good for ≤4B models, encoders and latency measurement on a known GPU, and is not a general escape hatch.
Key performance results
open-alternative-jev — Slightly better accuracy than Jev: 73.7% vs 72.7%, with better-calibrated probabilities.
AnyJev — Achieves similar accuracy to Jev without fine-tuning, and also produces well-calibrated confidence scores.
djev-dev — Performs very close to Jev on JevBench: ~73 vs ~74–75.
Nimble — A smaller model that still gets close to Jev: 90.1% vs 93.2% on its benchmark.
SemIf-OpenJev — Uses no fine-tuning and still ranks around #2–#4 on JevBench.
Open questions, and what the second post will do
Assumptions in this post. Every number is published by someone else, and star counts and rankings were read on 22–23 September 2026. JevBench's Jev score already differs between my two reads, 75.4 and 74.4, which tells you how fast this is moving.
What still looks open, as far as I can find:
I have not found a published run that puts the trained open models through one equal-calibration protocol on one frozen suite. If someone has done it since, that is the result I most want pointed at.
I have not seen several of these run on one Mac with one request shape. The published Mac numbers I could find span four different Macs: M1 Pro, M3 Max, M5 and M5 Pro.
"Does it know it doesn't know" is the least crowded question. Jev answers at ≥0.9 confidence on 9% of items whose deciding evidence was removed, kev-9b on 5%, kev-8b on 26%. scienthoon's unknowable-priority suite pins Jev at 44.7% accuracy while it states a mean probability of 0.74 on the level it chose.
How much of any gap is the readout rather than the model.
The second post: one machine, one frozen suite, the same calibration budget for every model, at least two option orders per choice item, latency reported in families that are never combined into one ranking, and raw responses published.
Roughly forty projects appeared in a week, and any ranking published today is stale in two, which is why the second post is a protocol rather than a scoreboard.
If I have mis-read a project or graded a number wrongly, I would like to know. Several of these authors document their own weaknesses better than most commercial vendors do, and they deserve to be quoted accurately.
Appendix A — sources
Every number above traces to one of these. All public, all read on 22–23 September 2026.
TypeSafe (class V)
Independent analyses of Jev (class T)
Archer Hume — Jev's architecture unmasked — black-box probing, tokenizer matching against 192 candidates, option-interaction experiments, MMLU ECE 0.0313
Aman Arora — a first look at Jev — primitives, context split, 299 ms against 1,114 ms
Arcturus Labs — Jev trades text generation for instant calibrated decisions — coin-flip calibration probe
anth.us — can you trust Jev confidence? — AUROC 0.83, isotonic ECE 0.008
The projects
laya · laya-mlx · kev · SemIf · GLiNER2 · NanoJev · nimble · von · decider · jeff · AnyJev · open-alternative-jev · djev-dev
Benchmarks, harnesses and datasets
JevBench — 534 frozen decisions, 66 ranked rows, hard tier hashed before any run
Decision Index — 132,422 requests over 37 benchmarks,
transformersengine withdevice=mpsjev-ood-calibration — a regenerable 900-ticket synthetic suite with an unknowable-by-design label
LocalLLaMA/typed-decisions— 400 cases × 5 typed questions, the de-facto community suiteLuni/laya-jev-benchmark— third-party scorer with a Jev 1.13.0 row
Discovery aids
systemonemodels.org/examples/alternatives — 20 listed reproductions, but it misreports kev, decider and SemIf as having no published metrics. Use it to find things, not to cite them.
GitHub
topic:system-oneandtopic:jev.
Unverified, and flagged as such
Reports of "$5 free credit, no waitlist" for the TypeSafe console are secondary and SEO-flavoured; typesafe.ai still renders a waitlist element.
nimble's repository has no LICENSE file although its model card says Apache-2.0. NanoJev's Hugging Face model-card licence is unset.
JevBench's Jev score differs between my two reads (75.4 and 74.4). Re-read before citing.
Appendix B — live additions
List of useful Jev related links as discovered.
| Added | Detail |
|---|---|
| 25 Sep | Jev replications, models, interpretations and papers in one place: |
| https://hanxiao.io/all-about-jev/#v=table&cat=model&sort=score&d=-1 | |

