<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Infocruncher]]></title><description><![CDATA[ML, AI, Maths, Software and the future.]]></description><link>https://infocruncher.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/6ab31cad7ceb28928bec518f/063f7faf-a6f1-4dd6-b902-7e762467121f.png</url><title>Infocruncher</title><link>https://infocruncher.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Sat, 03 Oct 2026 14:15:43 GMT</lastBuildDate><atom:link href="https://infocruncher.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Static embeddings: a practical guide to model2vec]]></title><description><![CDATA[Sentence-transformer embeddings work well, but they are slow on a CPU. For a few thousand texts, the speed doesn't matter. When you embed millions of texts, or need embeddings within a tight latency b]]></description><link>https://infocruncher.hashnode.dev/practical-guide-to-model2vec</link><guid isPermaLink="true">https://infocruncher.hashnode.dev/practical-guide-to-model2vec</guid><category><![CDATA[model2vec]]></category><category><![CDATA[#Embeddings]]></category><category><![CDATA[classification]]></category><dc:creator><![CDATA[Dylan Hogg]]></dc:creator><pubDate>Fri, 25 Sep 2026 02:30:12 GMT</pubDate><content:encoded><![CDATA[<p><a href="https://sbert.net">Sentence-transformer</a> embeddings work well, but they are slow on a CPU. For a few thousand texts, the speed doesn't matter. When you embed millions of texts, or need embeddings within a tight latency budget, the compute cost and latency add up quickly.</p>
<p><a href="https://github.com/MinishLab/model2vec">model2vec</a> reduces that cost. It distils a sentence-transformer into a static lookup table of token vectors. <a href="https://github.com/MinishLab">MinishLab</a>, the team behind it, report that the distilled model is up to 500x faster and up to 50x smaller than the original. The distilled model is also less accurate, and the size of the drop depends on the task. This post explains how model2vec works, which model to pick, how it compares with transformers and older approaches, and which tasks I think it suits.</p>
<h2>What model2vec is</h2>
<p>model2vec is an open-source library from MinishLab, first released in 2024. Its only major runtime dependency is numpy.</p>
<p>Distillation has three steps:</p>
<ol>
<li><p>Pass every token in the vocabulary through a teacher sentence-transformer. This gives one vector per token.</p>
</li>
<li><p>Reduce the vectors' dimensions with PCA (principal component analysis).</p>
</li>
<li><p>Down-weight common tokens using their estimated frequency (based on Zipf's law), so words like "the" contribute less to the final embedding.</p>
</li>
</ol>
<p>To embed a text, the model tokenises it, looks up each token's vector, and averages the vectors. There is no attention layer, so it runs fast on a CPU and doesn't need a GPU.</p>
<p>Because each token's vector is fixed, a token gets the same vector in every context. The model also ignores word order.</p>
<pre><code class="language-python">from model2vec import StaticModel

model = StaticModel.from_pretrained("minishlab/potion-base-8M")
embeddings = model.encode(["The cat sat on the mat", "A kitten rested on the rug"])
</code></pre>
<p>You can also distil your own model from any sentence-transformer. MinishLab say this takes about 30 seconds on a CPU and needs no training data. You can also add your own domain vocabulary. An optional second step, <a href="https://github.com/MinishLab/tokenlearn">Tokenlearn</a>, trains the static vectors to match the teacher's sentence embeddings over a corpus. The potion models in the next section were trained with Tokenlearn.</p>
<h2>Which model to pick</h2>
<table>
<thead>
<tr>
<th>Model</th>
<th>Params</th>
<th>Dims</th>
<th>Size</th>
<th>Use</th>
</tr>
</thead>
<tbody><tr>
<td>potion-base-8M</td>
<td>7.5M</td>
<td>256</td>
<td>30 MB</td>
<td>General English, small &amp; fast</td>
</tr>
<tr>
<td>potion-base-32M</td>
<td>32.3M</td>
<td>512</td>
<td>129 MB</td>
<td>Best general English quality</td>
</tr>
<tr>
<td>potion-retrieval-32M</td>
<td>32.3M</td>
<td>512</td>
<td>129 MB</td>
<td>English retrieval</td>
</tr>
<tr>
<td>potion-multilingual-128M</td>
<td>128M</td>
<td>256</td>
<td>512 MB</td>
<td>101 languages</td>
</tr>
<tr>
<td>potion-code-16M-v2</td>
<td>16.2M</td>
<td>256</td>
<td>32 MB</td>
<td>Code search</td>
</tr>
</tbody></table>
<p>The English models are distilled from <code>bge-base-en-v1.5</code>. The multilingual model is much larger because its vocabulary has about 500,000 tokens, and each token needs its own vector. The sizes are the model files on Hugging Face, stored as float32, except the code model, which is float16. MinishLab also publish smaller models: potion-base-2M (7.6 MB) and potion-base-4M (15 MB). The potion models replace the older <code>M2V_*</code> models.</p>
<p>I suggest starting with potion-base-8M, and moving to potion-base-32M if your evaluation shows a quality gap.</p>
<h2>How it compares with sentence-transformers</h2>
<p>MinishLab compare their models with <code>all-MiniLM-L6-v2</code>, which was released in 2021. Newer small models score much higher than MiniLM, so I've added two for comparison: <code>bge-small-en-v1.5</code> and <code>embeddinggemma-300m</code> (released September 2025). The scores below are from the official MTEB(eng, v2) results (MTEB is the <a href="https://huggingface.co/spaces/mteb/leaderboard">Massive Text Embedding Benchmark</a>), averaged over task types.</p>
<table>
<thead>
<tr>
<th>Task</th>
<th>potion-8M</th>
<th>potion-32M</th>
<th>MiniLM-L6</th>
<th>bge-small</th>
<th>EmbeddingGemma</th>
</tr>
</thead>
<tbody><tr>
<td>Params</td>
<td>7.5M</td>
<td>32M</td>
<td>23M</td>
<td>33M</td>
<td>308M</td>
</tr>
<tr>
<td><strong>Average</strong></td>
<td>51.08</td>
<td>52.13</td>
<td>55.93</td>
<td>59.93</td>
<td>65.11</td>
</tr>
<tr>
<td>Classification</td>
<td>70.34</td>
<td>71.70</td>
<td>69.25</td>
<td>76.56</td>
<td>87.55</td>
</tr>
<tr>
<td>Clustering</td>
<td>39.74</td>
<td>41.25</td>
<td>44.90</td>
<td>47.02</td>
<td>56.55</td>
</tr>
<tr>
<td>STS</td>
<td>72.91</td>
<td>73.93</td>
<td>78.95</td>
<td>81.28</td>
<td>83.61</td>
</tr>
<tr>
<td>Retrieval</td>
<td>31.11</td>
<td>32.67</td>
<td>42.92</td>
<td>53.86</td>
<td>55.69</td>
</tr>
</tbody></table>
<p>The table shows four things:</p>
<ul>
<li><p>On classification, potion matches or slightly beats MiniLM. bge-small scores about 5 points higher than potion-32M, and EmbeddingGemma about 16 points higher.</p>
</li>
<li><p>On clustering and STS (semantic textual similarity), potion-32M scores 4 to 17 points lower than the transformers.</p>
</li>
<li><p>On retrieval, potion-32M scores 10 to 23 points lower, which is the largest gap.</p>
</li>
<li><p>potion-base-32M keeps about 86% of its teacher's score (52.13 against 60.77 for <code>bge-base-en-v1.5</code>).</p>
</li>
</ul>
<p>model2vec scores lower than modern transformers on every task type. The reason to use it is speed and cost. MinishLab report up to 500x faster inference, and the smallest model is about 8 MB. The numpy-only runtime also gives faster start-up and smaller deployments.</p>
<p>Static embeddings can't represent meaning that depends on context. "Not good" ends up close to "good", word order has no effect, and a word like "bank" gets the same vector in every sentence.</p>
<p>model2vec and sentence-transformers are compatible. sentence-transformers can load model2vec models through its <code>StaticEmbedding</code> module. Tom Aarsen at Hugging Face has also used sentence-transformers to train related static models, such as <code>static-retrieval-mrl-en-v1</code>, with contrastive loss.</p>
<h2>How it compares with other approaches</h2>
<ul>
<li><p><strong>TF-IDF and BM25</strong> are the fastest option and a strong baseline. They only match exact words, so they treat "car" and "vehicle" as unrelated.</p>
</li>
<li><p><strong>GloVe and BPEmb</strong> are older static embeddings. Their MTEB averages are 45.82 and 41.74, against 51.08 for potion-8M. I think model2vec scores higher mainly because it uses subword tokens and learns from a modern teacher model, while GloVe learns from raw word co-occurrence counts.</p>
</li>
<li><p><a href="https://github.com/dleemiller/WordLlama"><strong>WordLlama</strong></a> also builds static token vectors, but takes them from an LLM's token embeddings.</p>
</li>
<li><p><strong>Small transformers</strong> such as MiniLM, bge-small and granite-embedding-small-english-r2 model context, but need far more compute per text. ONNX and int8 quantisation make them faster, but they are still much slower than a lookup table.</p>
</li>
<li><p><strong>Small LLM-based embedders</strong> such as EmbeddingGemma-300m and Qwen3-Embedding-0.6B give the best quality at small sizes, with MTEB averages around 65. They have 10 to 25 times more parameters than MiniLM, so they're expensive to run on a CPU at high volume.</p>
</li>
<li><p><a href="https://github.com/huggingface/setfit"><strong>SetFit</strong></a> is a strong few-shot classifier, but it runs a transformer. In MinishLab's classification benchmark, a fine-tuned model2vec classifier was about 35x faster than SetFit on a CPU.</p>
</li>
<li><p><strong>Embedding APIs</strong> give strong quality, but you pay per request, wait on network calls, and send your data to a third party.</p>
</li>
</ul>
<table>
<thead>
<tr>
<th>Approach</th>
<th>Speed on CPU</th>
<th>Quality</th>
<th>Uses context</th>
</tr>
</thead>
<tbody><tr>
<td>TF-IDF / BM25</td>
<td>Fastest</td>
<td>Low</td>
<td>No</td>
</tr>
<tr>
<td>GloVe / BPEmb</td>
<td>Very fast</td>
<td>Low</td>
<td>No</td>
</tr>
<tr>
<td>model2vec</td>
<td>Very fast</td>
<td>Medium</td>
<td>No</td>
</tr>
<tr>
<td>Small transformers</td>
<td>Slow</td>
<td>Good</td>
<td>Yes</td>
</tr>
<tr>
<td>LLM-based embedders</td>
<td>Very slow</td>
<td>Best (local)</td>
<td>Yes</td>
</tr>
<tr>
<td>Embedding APIs</td>
<td>Network-bound</td>
<td>Best</td>
<td>Yes</td>
</tr>
</tbody></table>
<h2>Use cases</h2>
<p><strong>Text classification.</strong> Encode your texts and train a simple classifier on the embeddings:</p>
<pre><code class="language-python">from sklearn.linear_model import LogisticRegression

clf = LogisticRegression(max_iter=1000).fit(model.encode(train_texts), train_labels)
preds = clf.predict(model.encode(test_texts))
</code></pre>
<p>For better accuracy, use <code>StaticModelForClassification</code> (in the <code>model2vec[train]</code> extra), which fine-tunes the embeddings and the classifier together. Accuracy drops most on tasks where the label depends on negation or word order, such as some sentiment tasks.</p>
<p><strong>Clustering.</strong> Encode the documents, then cluster the embeddings with KMeans or HDBSCAN. Embedding millions of documents is cheap enough to do on a laptop. The clusters will be less coherent than a transformer's: on MTEB clustering, the transformers above score 4 to 15 points higher than potion-32M.</p>
<pre><code class="language-python">from sklearn.cluster import KMeans

cluster_ids = KMeans(n_clusters=20).fit_predict(model.encode(docs))
</code></pre>
<p><strong>Semantic similarity.</strong> The potion models return normalised vectors, so the dot product of two embeddings is their cosine similarity. This works well for near-duplicate detection and finding related texts. Because the model ignores word order, "dog bites man" and "man bites dog" get identical vectors.</p>
<pre><code class="language-python">a, b = model.encode(["The film was great", "I loved the movie"])
similarity = a @ b
</code></pre>
<p><strong>Feature generation.</strong> Embeddings make good input features for gradient-boosted trees or other models. Encoding is cheap, so you can add a text column's embeddings to your existing tabular features without slowing the pipeline much:</p>
<pre><code class="language-python">import numpy as np

features = np.hstack([tabular_features, model.encode(df["description"].tolist())])
</code></pre>
<p><strong>Retrieval.</strong> This post doesn't cover retrieval in depth. potion-retrieval-32M scores 35.06 on MTEB retrieval, against 42.92 for MiniLM and 53.86 for bge-small. I think it works as a fast first stage in front of a reranker, but it can't replace a strong retrieval model.</p>
<h2>When not to use it</h2>
<p>Don't use model2vec when your task depends on negation, word order or long-range context, or when retrieval quality is the main goal. For other tasks, I start with model2vec, evaluate it on my own data, and switch to a small transformer only where the quality gap matters.</p>
<h2>Takeaways</h2>
<ul>
<li><p>Static embeddings trade some quality for a very large speed gain.</p>
</li>
<li><p>On classification, model2vec matches MiniLM, but modern small transformers score clearly higher.</p>
</li>
<li><p>If your workload is CPU-bound and high-volume, try model2vec first. It may be accurate enough for your task.</p>
</li>
</ul>
<h2>Sources</h2>
<ul>
<li><p><a href="https://github.com/MinishLab/model2vec">MinishLab/model2vec</a> README and <a href="https://github.com/MinishLab/model2vec/blob/main/results/README.md">results</a></p>
</li>
<li><p><a href="https://huggingface.co/minishlab">MinishLab models on Hugging Face</a></p>
</li>
<li><p><a href="https://github.com/embeddings-benchmark/results">MTEB results</a>, benchmark MTEB(eng, v2), averaged over task types</p>
</li>
<li><p><a href="https://huggingface.co/blog/static-embeddings">Train 400x faster static embedding models with Sentence Transformers</a> (Tom Aarsen, Hugging Face)</p>
</li>
</ul>
]]></content:encoded></item><item><title><![CDATA[Open alternatives to Jev]]></title><description><![CDATA[The open System One models you can run yourself
Jev is a week old and has been hard to avoid since launch. TypeSafe's "System One" model is built to be called by software. You post one state plus a ma]]></description><link>https://infocruncher.hashnode.dev/open-alternatives-to-jev</link><guid isPermaLink="true">https://infocruncher.hashnode.dev/open-alternatives-to-jev</guid><category><![CDATA[jev]]></category><category><![CDATA[survey]]></category><dc:creator><![CDATA[Dylan Hogg]]></dc:creator><pubDate>Wed, 23 Sep 2026 08:40:56 GMT</pubDate><content:encoded><![CDATA[<p><em>The open System One models you can run yourself</em></p>
<p><a href="https://typesafe.ai/">Jev</a> is a week old and has been hard to avoid since launch. TypeSafe's "System One" model is built to be called by software. You post one <code>state</code> plus a map of typed questions, and you get back typed values and calibrated probability distributions in a single forward pass. There are no generated tokens, nothing to parse, and no retry when the model decides to explain itself first. It caught on because it solves a problem most of us have been working around with constrained decoding, JSON schemas and regexes, because it costs $0.042 per million input tokens with output free, and because TypeSafe's docs are more candid about the model's failure modes than most vendors manage.</p>
<p>The interface is also easy to reimplement, and a lot of people have. This post catalogues the thirteen projects that clear the bar set out under "How I picked what to include". Roughly forty further candidates turned up in a sweep and did not clear it. Eleven of the thirteen run on an M1 Max with 64 GB of unified memory, which is the machine a follow-up post will use; NanoJev and djev-dev are documented rather than run, because leaving them out would misrepresent the field.</p>
<p>I summarise what each project does and where it matches or differs from Jev. I do not rank them: every number here was measured by someone else, and I say who each time. A second post will measure a subset on one machine with one suite, which is the only way I know to turn a catalogue into a comparison.</p>
<p>Interestingly, two of the most methodologically careful projects here, AnyJev and open-alternative-jev, train nothing at all. Both wrap a stock open model, read its option logits properly, and reach Jev's accuracy with better-calibrated probabilities. Both are single-benchmark, single-model-family results, so this settles nothing, but it does change what you should be asking of the nine trained open decision models.</p>
<p>→ <a href="https://infocruncher.hashnode.dev/open-alternatives-to-jev#whether-you-can-use-it">Jump straight to the results</a></p>
<h2>What is Jev?</h2>
<p>This section covers only the details the comparisons depend on. TypeSafe's <a href="https://docs.typesafe.ai/introduction/quickstart">own docs</a> and the four independent analyses in the appendix cover the rest.</p>
<p>There are three primitives (V):</p>
<table>
<thead>
<tr>
<th>Primitive</th>
<th>You supply</th>
<th>You get back</th>
</tr>
</thead>
<tbody><tr>
<td><code>noul</code></td>
<td>optional descriptions of what true and false mean</td>
<td><code>noul</code>: P(yes) as a float. <strong>No</strong> <code>confidence</code> <strong>field.</strong></td>
</tr>
<tr>
<td><code>choice</code></td>
<td>option → rubric description, <strong>max 255 options</strong></td>
<td><code>choice</code> (argmax), <code>probabilities</code>, <code>confidence</code></td>
</tr>
<tr>
<td><code>score</code></td>
<td>an ordered array of <strong>2–10</strong> level descriptions</td>
<td><code>score</code> (probability-weighted, lands between levels), <code>legend</code>, <code>probabilities</code>, <code>confidence</code></td>
</tr>
</tbody></table>
<p>Post-training is described as RLCD, reinforcement learning for calibrated decisions: an outcome-based objective against proper scoring rules rather than RLHF against human preference. The commercial envelope (V): model <code>jev-1.13.0</code>, $0.042 per million input tokens with output free, 64k context of which 32k is <code>state</code> plus the single longest question, text only, English first, and <strong>no fine-tuning or customisation of any kind</strong>. That last row matters more than it looks, because every open competitor's main lever is fine-tuning on your own labels.</p>
<p><code>confidence</code> is a peakedness statistic, not a probability of being correct. For a choice over K options it is <code>(K · p_max − 1) / (K − 1)</code>, normalised so a uniform distribution gives 0 and a one-hot gives 1. It says nothing about whether the argmax is right. TypeSafe say this themselves and return the full <code>probabilities</code> vector so you can compute your own, and two independent measurements suggest you should: anth.us finds AUROC 0.83 for ranking correct from incorrect but 7–15 points of overconfidence as a probability, and scienthoon finds the field "was never better than the max probability and sometimes much worse" (both T).</p>
<p>The vendor's own <a href="https://docs.typesafe.ai/model-jaggedness/jev-1.13">jaggedness page</a> lists nine failure modes: literal reading of implied intent, arithmetic and counting, date comparison, indirection, large irrelevant state acting as a distractor, injected content, contradictory instructions, no guaranteed invariants between related questions, and generation. Their guidance — move arithmetic into code, filter the state, compare dates in code — amounts to saying Jev is a semantic verifier rather than a reasoner. I think that is a fair description and a reasonable thing to sell.</p>
<h2>Why you would look for alternatives to Jev?</h2>
<p>Cost is not the reason. At $0.042/Mtok input with output free, and JevBench's measured ~950 input tokens per decision (V + T), that is roughly <strong>4 cents per 1,000 decisions</strong>. scienthoon's entire 4,621-call study cost about 6 cents. Check the arithmetic before you self-host to save money.</p>
<p>Four reasons that do hold up:</p>
<ol>
<li><p><strong>Trainability.</strong> Jev offers no fine-tuning, no LoRA and no per-account weights (V). You adapt through <code>state</code> and <code>criteria</code> only. Nine of the thirteen projects below take your labels, and nimble's jump from 66% to 90% on 2,676 examples shows how cheaply that can pay.</p>
</li>
<li><p><strong>Latency.</strong> Hosted Jev is around 250 ms wall clock from a US or European origin, consistent across three independent measurements, against roughly 160 ms of upstream service time. A local encoder is 13–130 ms with no network involved.</p>
</li>
<li><p><strong>Data residency.</strong> Zero data retention is an enterprise option. Running locally keeps the data on your own hardware.</p>
</li>
<li><p><strong>Pinning and continuity.</strong> Aliases move, the docs say rate limits can change without notice (V), and the Vercel AI Gateway route cannot pin a version at all. That matters once you have tuned thresholds against one build.</p>
</li>
</ol>
<p>The counterweight is that Jev is still ahead on knowledge-heavy decisions, which is where decider and kev both report their largest gaps, and that the alternatives were all days old when I wrote this.</p>
<h2>How I picked what to include</h2>
<p>A sweep of GitHub <code>topic:system-one</code> and <code>topic:jev</code>, a full-text repository search and a third-party index turned up roughly 40 projects beyond the obvious ones, plus about 20 "awesome-jev" link collections. Four criteria, all of which must hold:</p>
<ol>
<li><p>It implements the System One contract: typed primitives with a probability distribution per question, read from the model rather than generated as text. A service that calls the hosted Jev API is an application, not a comparator.</p>
</li>
<li><p>It runs locally and reproducibly, with open weights or a stock open-weight base plus published inference code.</p>
</li>
<li><p>It is pinned and licensed.</p>
</li>
<li><p>It has been measured, or ships a bench that can measure it.</p>
</li>
</ol>
<p>One warning about discovery aids: the widely-linked index at <a href="https://systemonemodels.org/examples/alternatives/">systemonemodels.org</a> lists kev, decider and SemIf as having no published metrics, which is wrong for all three. Use it to find things, not to cite them.</p>
<h3>Evidence grading</h3>
<p>I grade every number and don't mix classes inside a comparison:</p>
<table>
<thead>
<tr>
<th>Class</th>
<th>Meaning</th>
</tr>
</thead>
<tbody><tr>
<td><strong>V</strong></td>
<td>Verified spec — read from vendor docs, repo code or manifests</td>
</tr>
<tr>
<td><strong>T</strong></td>
<td>Third-party measurement — someone independent ran it and published raw data</td>
</tr>
<tr>
<td><strong>S</strong></td>
<td>Self-reported — the authors measured their own model against a number they did not produce</td>
</tr>
</tbody></table>
<p>Almost every "X vs Jev" table in these repos compares a class S number against a class T or S number, measured on different items, prompts, dates and hardware. Several of the repos say so themselves. Pooling those numbers into one ranking would produce a ranking that is confidently wrong.</p>
<h3>What was excluded, and why</h3>
<table>
<thead>
<tr>
<th>Excluded</th>
<th>Reason</th>
</tr>
</thead>
<tbody><tr>
<td><a href="https://github.com/ekzhang/openjev-sglang">openjev-sglang</a></td>
<td>No licence file (criterion 3). Also 35B</td>
</tr>
<tr>
<td><a href="https://github.com/Heman10x-NGU/openJev-verdict-2.0">openJev-verdict-2.0</a></td>
<td>Verdict 2.0's weights are unobtainable (criterion 2)</td>
</tr>
<tr>
<td><a href="https://github.com/vinnylarouge/jevlike">jevlike</a></td>
<td>No accuracy or calibration metric (criterion 4). A teaching repo, and a good one</td>
</tr>
<tr>
<td><a href="https://github.com/OmniJev/PlayJev">PlayJev</a>, <a href="https://github.com/PsiACE/dohnuts">dohnuts</a>, jev-visual</td>
<td>Multimodal or GUI-game specialists (criterion 1 in spirit)</td>
</tr>
<tr>
<td><a href="https://github.com/9pings/notjev">notjev</a>, SiliconLabAI/OpenJev, iamaamir/system-one</td>
<td>Provider-neutral wrappers over someone else's endpoint (criterion 1)</td>
</tr>
<tr>
<td>~13 small repos: mini-jev, litjev, jevmlx, fastjev, minojev, jevfire and others</td>
<td>Created 16–22 September, no published metric and no runnable bench (criterion 4). Several look plausible; none is yet measurable</td>
</tr>
</tbody></table>
<h2>The alternative field</h2>
<p>The first table says what a project <em>is</em>; the second says whether you can <em>use</em> it. Row order is the same in both.</p>
<h3>What each alternative is</h3>
<table>
<thead>
<tr>
<th>Project</th>
<th>Created</th>
<th>Architecture</th>
<th>Params</th>
<th>Licence</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Jev 1.13.0</strong> (reference)</td>
<td>—</td>
<td>inferred sparse MoE decoder + linear head</td>
<td>~10B active (inferred, T)</td>
<td>commercial, hosted only</td>
</tr>
<tr>
<td><a href="https://github.com/NandhaKishorM/laya">laya</a></td>
<td>18 Sep</td>
<td>ModernBERT-large / mmBERT-base encoders, three checkpoints + router</td>
<td>421M / 322M / 421M</td>
<td>Apache-2.0</td>
</tr>
<tr>
<td><a href="https://github.com/mizorewww/laya-mlx">laya-mlx</a></td>
<td>19 Sep</td>
<td>MLX port of the three laya checkpoints</td>
<td>same</td>
<td>Apache-2.0</td>
</tr>
<tr>
<td><a href="https://github.com/jaredpalmer/kev">kev</a></td>
<td>17 Sep</td>
<td>Qwen3.5-Base + LoRA + pointer head</td>
<td>0.8B / 4B / 9B</td>
<td>Apache-2.0</td>
</tr>
<tr>
<td><a href="https://github.com/TheoLeeCJ/SemIf-OpenJev">SemIf</a></td>
<td>16 Sep</td>
<td>stock Qwen3.5-4B read through option logits, <strong>no training</strong></td>
<td>4B</td>
<td>MIT (code)</td>
</tr>
<tr>
<td><a href="https://github.com/fastino-ai/GLiNER2">GLiNER2</a></td>
<td>Jul 2025</td>
<td>encoder, schema-driven extraction and classification</td>
<td>~300M</td>
<td>Apache-2.0</td>
</tr>
<tr>
<td><a href="https://github.com/TianyuCodings/NanoJev">NanoJev</a></td>
<td>17 Sep</td>
<td>Qwen3-0.6B + decision heads</td>
<td>0.6B</td>
<td>MIT (repo); HF repo unset</td>
</tr>
<tr>
<td><a href="https://github.com/bespokelabsai/nimble">nimble</a></td>
<td>18 Sep</td>
<td>Qwen3.5-9B + LoRA r16</td>
<td>9B</td>
<td>card says Apache-2.0; <strong>repo has no LICENSE file</strong></td>
</tr>
<tr>
<td><a href="https://github.com/wfzyx/von">von</a></td>
<td>18 Sep</td>
<td>ModernBERT-Large bidirectional encoder</td>
<td>395M</td>
<td>Apache-2.0</td>
</tr>
<tr>
<td><a href="https://github.com/Mapika/decider">decider</a></td>
<td>16 Sep</td>
<td>Qwen3.5-Base + decision head</td>
<td>1.9B / 34.7B-A3B (3B active)</td>
<td>Apache-2.0</td>
</tr>
<tr>
<td><a href="https://github.com/logan-markewich/jeff">jeff</a></td>
<td>19 Sep</td>
<td>GLiFormer-large-v1 encoder</td>
<td>400M</td>
<td>MIT</td>
</tr>
<tr>
<td><a href="https://github.com/nokia-applied-research/AnyJev">AnyJev</a></td>
<td>21 Sep</td>
<td>readout layer over any LLM, <strong>no training</strong></td>
<td>any</td>
<td>Apache-2.0</td>
</tr>
<tr>
<td><a href="https://github.com/ikermoel/open-alternative-jev">open-alternative-jev</a></td>
<td>18 Sep</td>
<td>packed-state readout over any open LLM, <strong>no training</strong></td>
<td>any</td>
<td>Apache-2.0</td>
</tr>
<tr>
<td><a href="https://github.com/Davipar/djev-dev">djev-dev</a></td>
<td>Sep</td>
<td>DiffusionGemma-26B-A4B via vLLM</td>
<td>26B MoE, 4B active</td>
<td>Apache-2.0</td>
</tr>
</tbody></table>
<p>Star counts are a poor guide as of 23 Sep 2026. laya gained 2,583 stars in the 24 hours between my two research passes, and the two most methodologically careful projects in the list have 68 and 51 stars. Sort by something else.</p>
<h3>Whether you can use it</h3>
<p>These columns decide whether you can adopt a project. They are currently scattered across thirteen READMEs, and collecting them is useful.</p>
<table>
<thead>
<tr>
<th>Project</th>
<th>Primitives</th>
<th>Max choice options</th>
<th>Context</th>
<th>Drop-in <code>/v1/systemone</code>?</th>
<th>Apple Silicon</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Jev 1.13.0</strong> (reference)</td>
<td><code>noul</code>, <code>choice</code>, <code>score</code></td>
<td><strong>255</strong></td>
<td>64k (32k state + longest question)</td>
<td>it <em>is</em> the spec</td>
<td>hosted only</td>
</tr>
<tr>
<td><a href="https://github.com/NandhaKishorM/laya">laya</a></td>
<td>all three</td>
<td>no hard cap, degrades badly above ~20</td>
<td><strong>512</strong> English, <strong>1,024</strong> multilingual (state shares it with the options)</td>
<td>no — own Python API</td>
<td>CPU/MPS, 193–464 ms CPU</td>
</tr>
<tr>
<td><a href="https://github.com/mizorewww/laya-mlx">laya-mlx</a></td>
<td>all three</td>
<td>as laya</td>
<td>as laya, per checkpoint</td>
<td>no — own Python API</td>
<td>native MLX, <strong>13.4 ms P50</strong> (M3 Max)</td>
</tr>
<tr>
<td><a href="https://github.com/jaredpalmer/kev">kev</a></td>
<td>all three</td>
<td><strong>1–255</strong></td>
<td><strong>8,192</strong> serving (state + one question); trained at ≤384 state tokens</td>
<td><strong>yes, native</strong></td>
<td>yes, but slow on the Qwen3.5 generation</td>
</tr>
<tr>
<td><a href="https://github.com/TheoLeeCJ/SemIf-OpenJev">SemIf</a></td>
<td>choice only — a yes/no is two options</td>
<td>not published</td>
<td>base model's (Qwen3.5-4B)</td>
<td>no — own scorer</td>
<td>yes: MLX, MPS and llama.cpp</td>
</tr>
<tr>
<td><a href="https://github.com/fastino-ai/GLiNER2">GLiNER2</a></td>
<td>n/a — extraction schema</td>
<td>not published</td>
<td>encoder window; long documents go through a 384-word chunker</td>
<td>no</td>
<td>CPU-friendly</td>
</tr>
<tr>
<td><a href="https://github.com/TianyuCodings/NanoJev">NanoJev</a></td>
<td>all three</td>
<td><strong>2–255</strong> candidates</td>
<td>not published</td>
<td>no — own <code>/api/evaluate</code></td>
<td><strong>no</strong> — CUDA + ViZDoom</td>
</tr>
<tr>
<td><a href="https://github.com/bespokelabsai/nimble">nimble</a></td>
<td>all three</td>
<td><strong>≤26</strong> (A–Z encoding)</td>
<td><strong>2,048-token prompt cap</strong></td>
<td>no — own <code>ParallelScorer</code></td>
<td>yes, MLX first-class</td>
</tr>
<tr>
<td><a href="https://github.com/wfzyx/von">von</a></td>
<td>all three</td>
<td>not published — <code>two_stage_choice</code> exists for taxonomies above ~25</td>
<td>not published</td>
<td><strong>yes, native</strong></td>
<td>yes, MPS</td>
</tr>
<tr>
<td><a href="https://github.com/Mapika/decider">decider</a></td>
<td>all three</td>
<td><strong>2–255</strong></td>
<td><strong>32k</strong></td>
<td><strong>yes, native</strong></td>
<td>yes, MPS merged 22 Sep</td>
</tr>
<tr>
<td><a href="https://github.com/logan-markewich/jeff">jeff</a></td>
<td>all three</td>
<td><strong>64</strong> (<code>JEFF_MAX_LABELS</code>, configurable)</td>
<td><strong>20,000 state characters</strong> (<code>JEFF_MAX_STATE_CHARS</code>); no token figure published</td>
<td><strong>yes, native</strong></td>
<td>yes — CUDA → MPS → CPU</td>
</tr>
<tr>
<td><a href="https://github.com/nokia-applied-research/AnyJev">AnyJev</a></td>
<td>all three via readout</td>
<td><strong>26</strong></td>
<td>base model's</td>
<td>not yet — on the roadmap</td>
<td>via transformers/MPS</td>
</tr>
<tr>
<td><a href="https://github.com/ikermoel/open-alternative-jev">open-alternative-jev</a></td>
<td>all three — <code>Choice</code>, <code>yes_no</code>, <code>scale</code></td>
<td><strong>26</strong> (A–Z, each letter one token)</td>
<td>base model's</td>
<td>no — own packed API</td>
<td>via transformers</td>
</tr>
<tr>
<td><a href="https://github.com/Davipar/djev-dev">djev-dev</a></td>
<td>all three, plus images as options</td>
<td>not published</td>
<td>not published</td>
<td>no — Jev-shaped body at its own <code>/v1/request</code></td>
<td><strong>no</strong> — CUDA-only vLLM</td>
</tr>
</tbody></table>
<p>"Not published" means I went through the repo and the figure genuinely isn't there. Six of the thirteen don't state an option cap and five don't state a context limit, and those are exactly the numbers that decide whether you can adopt something. Most of these repos publish an accuracy number against Jev before they publish their limits.</p>
<p>Two things stand out:</p>
<ul>
<li><p>kev, von, decider and jeff implement TypeSafe's actual <code>/v1/systemone</code> wire format, so the official SDK points at them with an environment variable change. For adoption that matters more than any accuracy number. kev's CI runs TypeSafe's own example requests against its server.</p>
</li>
<li><p><strong>Check the option cap before you read the accuracy table.</strong> A 40-option routing problem rules out nimble, AnyJev and open-alternative-jev immediately, degrades laya, and is over jeff's default limit, whatever those projects score.</p>
</li>
</ul>
<h2>The projects</h2>
<p>Each project follows the same template, and carries the repo link plus the exact commit and checkpoint revisions I read, because in a field this young an unpinned claim stops being checkable within days. Those pins are what I read, not what I recommend you run.</p>
<h3>1. laya</h3>
<p><a href="https://github.com/NandhaKishorM/laya">github.com/NandhaKishorM/laya</a></p>
<p><em>Read 23 Sep 2026 at</em> <a href="https://github.com/NandhaKishorM/laya/commit/725e020adba4f17ab838b2561a56fd20e05b022c"><code>725e020</code></a> <em>· weights</em> <a href="https://huggingface.co/convaiinnovations/laya/tree/1c5edc17a7acd8701df6fc341c0d179f1c62c982"><code>convaiinnovations/laya@1c5edc1</code></a><em>,</em> <a href="https://huggingface.co/convaiinnovations/laya-typed-decisions/tree/f9ab0b228f0fc0f14d873dbc99038f135c2da1b2"><code>convaiinnovations/laya-typed-decisions@f9ab0b2</code></a></p>
<p>Three encoder checkpoints — ModernBERT-large 421M, mmBERT-base 322M, and a typed-decisions fine-tune — with a script-detecting router that picks between them in under 0.5 ms. 2.4 GB total, Apache-2.0.</p>
<p>laya documents its own limits more openly than most of the field does. The base checkpoints score <strong>near chance zero-shot</strong> on its own typed-decisions benchmark: 0.362 and 0.342, against a 0.318 random baseline and a 0.461 majority-class floor. The headline 0.766 comes from the checkpoint fine-tuned on that benchmark's own training split, and the authors say so (S).</p>
<p>Three other limits. Raw ECE is 0.466 and falls to 0.081 only after per-(type, option-count) temperature fitting, and <code>laya-multilingual</code> ships with no fitted temperatures. High-cardinality choice collapses — Banking77 0.425 against Jev's 0.870 — because ~77 labels share a 192–256 token head budget, about three to four tokens per label. Khmer scores <strong>0.000 accuracy at 0.952 confidence</strong>, which is why the router exists.</p>
<p><strong>Differs from Jev:</strong> encoder rather than decoder, multilingual, and designed to be specialised rather than used zero-shot. The authors call it "a fast base to specialise, not a zero-shot decision engine", which is the fairest self-description in this field.</p>
<p><strong>Runs on:</strong> Mac (CPU/MPS) at 193–464 ms on CPU, 32.8 ms on a T4. It ships a Kaggle 2×T4 fine-tuning notebook.</p>
<h3>2. laya-mlx</h3>
<p><a href="https://github.com/mizorewww/laya-mlx">github.com/mizorewww/laya-mlx</a></p>
<p><em>Read 23 Sep 2026 at</em> <a href="https://github.com/mizorewww/laya-mlx/commit/0a859518634112655cb97c745dbf04f5191aaf13"><code>0a85951</code></a> <em>· weights</em> <a href="https://huggingface.co/aac6fef/laya-mlx/tree/20aed815fc6acde75733882e7ec0e3f28aeb9717"><code>aac6fef/laya-mlx@20aed81</code></a></p>
<p>An independent MLX port of all three laya checkpoints — encoder, decision transformer, scoring head and action head — with no PyTorch and no transformers runtime.</p>
<p>On an M3 Max in FP16, end to end including tokenisation: <strong>13.42 ms P50</strong> for one short question, 7.39 ms for the multilingual checkpoint, 146.8 and 395.0 questions per second at batch 50, under 1 GiB peak allocation (S, but unusually well documented).</p>
<p>Port fidelity is verified rather than asserted. All three checkpoints matched the upstream selected answer on 63/63 validation questions in both FP32 and FP16 — 378/378 comparisons — with pinned revisions, recorded weight hashes and 36 checksum-verified files. Almost no other port in this field documents fidelity at all.</p>
<p><strong>Differs from Jev:</strong> it inherits laya's architecture and laya's weakness. It sets the latency floor for local decisions, but those decisions come from checkpoints that score near chance zero-shot, so latency is not the constraint that matters here.</p>
<p><strong>Runs on:</strong> Apple Silicon only, macOS 14+.</p>
<h3>3. kev</h3>
<p><a href="https://github.com/jaredpalmer/kev">github.com/jaredpalmer/kev</a></p>
<p><em>Read 23 Sep 2026 at</em> <a href="https://github.com/jaredpalmer/kev/commit/557598fced1dada75dfbf36ed144dce309ac6ceb"><code>557598f</code></a> <em>· weights</em> <a href="https://huggingface.co/jaredpalmer/kev-4b/tree/485ace8703592fcf405488b262449990824cfed1"><code>jaredpalmer/kev-4b@485ace8</code></a><em>,</em> <a href="https://huggingface.co/jaredpalmer/kev-0.8b/tree/54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8"><code>jaredpalmer/kev-0.8b@54f4f87</code></a></p>
<p>Qwen3.5-Base plus LoRA plus a pointer head that scores each option's <code>&lt;/opt&gt;</code> hidden state against the question's <code>&lt;decide&gt;</code> hidden state. Three sizes: 0.8B, 4B and 9B.</p>
<p>Its harness, more than its model, makes it the most evaluation-ready project here: a genuine TypeSafe-compatible <code>/v1/systemone</code>, a benchmark runner that scores any System One endpoint over <code>--remote</code>, a live-Jev runner, and paired bootstrap confidence intervals between two saved runs. Its <code>evals/external/</code> already holds two other projects' test sets in this format with their published live-Jev results.</p>
<p>It also reports a metric I did not see elsewhere: <strong>share of decisions automatable at a 5% error budget, 0.45–0.57 against Jev's 0.70</strong> (S, but paired and on frozen items). Confident errors are 4.0% against Jev's 3.7%, after a fitted temperature of 2.1–2.4 baked into the checkpoint. Knowledge is where it loses, and the base model sets that gap: MMLU 0.74 against 0.90.</p>
<p><strong>Weakness that shapes the local shortlist:</strong> there are no fast Metal kernels for Qwen3.5's Gated DeltaNet layers, so PyTorch runs reference code. Median on an M5, five questions, ~230-token state: 329 ms for 0.8B, 779 ms for 4B, about 2 s for 9B. The previous Qwen3 generation is three to six times faster on the same Mac (4B at 174 ms). An MLX backend is the next planned change.</p>
<p><strong>Runs on:</strong> Mac, slowly, and any CUDA box.</p>
<h3>4. SemIf (formerly OpenJev)</h3>
<p><a href="https://github.com/TheoLeeCJ/SemIf-OpenJev">github.com/TheoLeeCJ/SemIf-OpenJev</a></p>
<p><em>Read 23 Sep 2026 at</em> <a href="https://github.com/TheoLeeCJ/SemIf-OpenJev/commit/1f2dea3e25379f9dfc98cb83c324f00ab5deda37"><code>1f2dea3</code></a> <em>· weights</em> <a href="https://huggingface.co/Qwen/Qwen3.5-4B/tree/851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a"><code>Qwen/Qwen3.5-4B@851bf6e</code></a></p>
<p><strong>Stock Qwen3.5-4B with no training at all</strong>, read through option logits. MLX, PyTorch/MPS and llama.cpp backends, all added in the last week.</p>
<p>It ranks around #2 to #4 on <a href="https://github.com/fstandhartinger/jevbench">JevBench</a> depending on the revision you read (T). On one RTX 3090, same model and same state, 21 binary criteria took 1.023 s through direct typed logits against 5.332 s for the autoregressive JSON equivalent, a 5.2× difference (S). The authors note that the two readouts agreed on only 18 of 21 criteria, so this is a systems comparison rather than a claim of semantic equivalence.</p>
<p><strong>Differs from Jev:</strong> nothing was trained. That makes SemIf the control condition this field needs. It replaces "which fine-tune is best" with the more useful question of whether any of this training beats reading a stock model properly.</p>
<p><strong>Runs on:</strong> Mac natively, three backends, ~8 GB in bf16.</p>
<h3>5. GLiNER2</h3>
<p><a href="https://github.com/fastino-ai/GLiNER2">github.com/fastino-ai/GLiNER2</a></p>
<p><em>Read 23 Sep 2026 at</em> <a href="https://github.com/fastino-ai/GLiNER2/commit/4abb6132c9efe737712003669dd1f569f6a5cb82"><code>4abb613</code></a> <em>· weights</em> <a href="https://huggingface.co/fastino/gliner2.5-base-v1/tree/1a8bc24e00dc7300b9017c81d63e3dcdabb26596"><code>fastino/gliner2.5-base-v1@1a8bc24</code></a></p>
<p>The one pre-Jev entry (July 2025): a ~300M encoder for structured extraction and classification. It was pulled into this conversation because it was already producing schema-driven typed outputs from an encoder.</p>
<p>Unlike every other project here, GLiNER2 has no published System One-framed comparison against Jev. I am listing it because it belongs in the lineage, not because there is a number to cite.</p>
<p><strong>Differs from Jev:</strong> extraction-first rather than decision-first, and no calibrated probability contract over a typed question set.</p>
<p><strong>Runs on:</strong> anything, including CPU.</p>
<h3>6. NanoJev</h3>
<p><a href="https://github.com/TianyuCodings/NanoJev">github.com/TianyuCodings/NanoJev</a></p>
<p><em>Read 23 Sep 2026 at</em> <a href="https://github.com/TianyuCodings/NanoJev/commit/76fdfc9ecdca45a9bcef17991a07d3041a87685a"><code>76fdfc9</code></a> <em>· weights</em> <a href="https://huggingface.co/C-Tianyu/NanoJev/tree/047b927b30882a1138fc504821b82ac145a4b81a"><code>C-Tianyu/NanoJev@047b927</code></a></p>
<p>Qwen3-0.6B with decision heads, trained on 18,760 decision questions across four games: ViZDoom Basic, ViZDoom Predict Position, a 50×50 maze, and Snake.</p>
<p>It beats Jev decisively on those games — ViZDoom Basic 128/128 against Jev's 56/128, maze solved in 225 attempts against 2,738 (S) — because it is a domain-specialist embodied-control model, not a general text-decision model. Benchmarking it against text decisions would be as misleading as testing Jev on Snake. Its sibling project <a href="https://github.com/TianyuCodings/JevHarness">JevHarness</a>, which generates LLM-authored task-specific Jev harnesses, is arguably more relevant to using Jev well.</p>
<p><strong>Runs on:</strong> CUDA only, 21.6 GB of checkpoints, plus ViZDoom.</p>
<h3>7. nimble (Bespoke Labs)</h3>
<p><a href="https://github.com/bespokelabsai/nimble">github.com/bespokelabsai/nimble</a></p>
<p><em>Read 23 Sep 2026 at</em> <a href="https://github.com/bespokelabsai/nimble/commit/38edc3b576f13179df785d412621d5cb1128d02d"><code>38edc3b</code></a> <em>· weights</em> <a href="https://huggingface.co/bespokelabs/Bespoke-Nimble-9B/tree/e93fabce8fcee46e7e8b45e2d955b8a4f0035917"><code>bespokelabs/Bespoke-Nimble-9B@e93fabc</code></a></p>
<p>Qwen3.5-9B plus LoRA r16, built in a day, trained on <strong>2,676 contrastive examples</strong>. The curation method is worth stealing: write two near-identical examples that differ in one fact that flips the label, verify both with separate model calls, then verify that removing either evidence sentence makes the fact unknowable.</p>
<p>Holdout results on 324 items: Jev 93.21%, Nimble-9B 90.12%, Qwen3.8-27B 84.88%, base Qwen3.5-9B 66.36% (S). That jump from 66% to 90% on 2,676 examples is the strongest single argument in the field that a small local model closes most of the gap cheaply.</p>
<p>The authors state their constraints plainly: a 2,048-token prompt cap, ≤26 enum options through A–Z encoding, and probabilities that are <strong>explicitly uncalibrated</strong> at temperature 1.0. The labels are synthetic and, in their words, no person has reviewed them.</p>
<p>There is also a licence discrepancy. The model card says Apache-2.0, but the repository has no LICENSE file. Resolve that before you depend on it.</p>
<p><strong>Runs on:</strong> Mac with 64 GB — the README says so explicitly — at 444 ms median on an M5 Pro, 106 ms on an H100. ~18 GB base plus a 0.17 GB LoRA merged on CPU.</p>
<h3>8. von</h3>
<p><a href="https://github.com/wfzyx/von">github.com/wfzyx/von</a></p>
<p><em>Read 23 Sep 2026 at</em> <a href="https://github.com/wfzyx/von/commit/14be898e7d4712423747380b9c1901e977743e3c"><code>14be898</code></a> <em>· weights</em> <a href="https://huggingface.co/wfzyx/von/tree/d8bb5e0745d8ee1fb65d536d6d4892d54d5a93fd"><code>wfzyx/von@d8bb5e0</code></a></p>
<p>ModernBERT-Large 395M bidirectional encoder, dual cross-entropy and Brier loss, 250k NLI examples plus ~290k multi-domain, a fitted temperature of 1.1692, the TypeSafe wire format, Python and TypeScript SDKs, and a claimed ~18 ms GPU latency at a 3.2 GB download.</p>
<p>The README's headline and its own table disagree. The bullet list claims "SOTA Empirical Accuracy: 91.23% ... surpassing published commercial alternatives". The benchmark table two paragraphs below shows <strong>von at 72.0% macro against Jev's 96.6%</strong> on the same 49-task suite, a 24-point deficit. The ViZDoom row (von 9.00 kills against Jev's 5.62) is real, but it is one narrow zero-shot control task. The document does not support its own framing.</p>
<p>von is architecturally interesting, cheap to run and plausibly among the fastest local options, but every published claim needs independent verification before anyone cites it.</p>
<p><strong>Runs on:</strong> Mac (MPS), CUDA, ROCm, CPU. Python ≥3.12.</p>
<h3>9. decider</h3>
<p><a href="https://github.com/Mapika/decider">github.com/Mapika/decider</a></p>
<p><em>Read 23 Sep 2026 at</em> <a href="https://github.com/Mapika/decider/commit/7557fe058e0b31ffcf754b7c5bb31a452454ce89"><code>7557fe0</code></a> <em>· weights</em> <a href="https://huggingface.co/Mapika/decider-2b/tree/fa996cea58e1c1d8d1ab4d7124154f303b017f95"><code>Mapika/decider-2b@fa996ce</code></a></p>
<p>Two released sizes on Qwen3.5-Base, 1.9B and 34.7B-A3B, plus a 0.8B and a vision variant. 32k context, 2–255 choice options, 2–10 score levels: it matches Jev's declared envelope more closely than anything else here (V).</p>
<p>It reports third-party rankings rather than its own (T). On JevBench: Jev #1 at 75.4, decider-35b-a3b #10 at 68.9, decider-2b #21 at 64.6. On <a href="https://github.com/apolinario/decision-index">Decision Index</a> v0.1, decider-35b-a3b is #4 of 32 and the highest-scoring <em>trained</em> model; the two entries above it are zero-training wrappers on stock 27B-class models.</p>
<p>It names where it loses: knowledge 0.51 against Jev's 0.69 per-area, and decider-2b's top-label ECE on JevBench hard items is 0.30, meaning it is confident where it is wrong. Its "Limits, stated plainly" section is the most candid in this list. One line from it: "rules written into the question are not followed at this size" — a one-sentence question scores 0.67, a paragraph of rules 0.24.</p>
<p>Apple Silicon is measured rather than asserted, and this is the only project in the list that did so. On an M1 Pro in float16, median request is 133 ms with the MPS patch against 171 ms without, and the MPS path scores 0.7553 accuracy / 0.0438 ECE against the published bf16 row's 0.756 / 0.041 on 1,500 held-out examples. That is a numerical-parity check on exactly this hardware class.</p>
<p><strong>Runs on:</strong> <code>pip install decider-ai</code>, <code>scripts/serve.sh</code> gives <code>/v1/systemone</code>. 1.9B is 3.8 GB; 34.7B is 69 GB in bf16 and will not fit in 64 GB unified memory.</p>
<h3>10. jeff</h3>
<p><a href="https://github.com/logan-markewich/jeff">github.com/logan-markewich/jeff</a></p>
<p><em>Read 23 Sep 2026 at</em> <a href="https://github.com/logan-markewich/jeff/commit/34b32f99a727c47b679adde33f4702a001e02979"><code>34b32f9</code></a> <em>· weights</em> <a href="https://huggingface.co/knowledgator/gliformer-large-v1/tree/d0a4e53d09cebe6bc963dd9be319d4279084bb2d"><code>knowledgator/gliformer-large-v1@d0a4e53</code></a></p>
<p>A self-hosted drop-in for the Jev API on GLiFormer-large-v1 (400M). Device order is CUDA → MPS → CPU, so it works on a Mac unmodified, and the official SDK points at it with one environment variable. All three primitives.</p>
<p>It publishes where it loses. Measured on 1,600 labelled items across eight datasets: AG News 75.5% against Jev's 90.5%. JevBench v1.2.2 score 66.9, #9 of 18, against Jev's 75.3 — and the author notes that this rank comes from cost, because jeff is #14 of 18 on intelligence (T). Latency is 151 ms p50 on an L4 via Modal against Jev's 129 ms.</p>
<p>A small encoder that is six times cheaper (~$2.6 against ~$15.6 per million single-question requests) and fifteen accuracy points worse is a real point on the trade-off curve, and no other project in this list occupies it.</p>
<p><strong>Runs on:</strong> Mac unmodified, ~0.8 GB.</p>
<h3>11. AnyJev (Nokia Applied Research)</h3>
<p><a href="https://github.com/nokia-applied-research/AnyJev">github.com/nokia-applied-research/AnyJev</a></p>
<p><em>Read 23 Sep 2026 at</em> <a href="https://github.com/nokia-applied-research/AnyJev/commit/54f1b3533639f52a406941706a4fdf08b2193589"><code>54f1b35</code></a> <em>· no weights of its own</em></p>
<p>AnyJev is a readout layer over any LLM rather than a model, and it trains nothing. Its thesis is that the problem is the readout, not the weights: <code>max_tokens=1</code> plus logprobs already gives you a ranking, but one that moves when you reorder the options, and a confidence number you cannot threshold on. Both are properties of the readout, and both are fixable without a single label.</p>
<p>There are three levels, and every returned decision carries its level so downstream code can refuse to act on the wrong one. <code>raw</code> is a restricted softmax over label tokens. <code>L0</code> adds cyclic-shift marginalisation — show K options in K rotations, combine in log space — plus prior correction, <strong>with no labels</strong>. <code>L1</code> adds temperature scaling on 100–500 labels.</p>
<p>Qwen3-8B on BANKING77 20-way, n=300 (S):</p>
<table>
<thead>
<tr>
<th></th>
<th>raw</th>
<th>L0</th>
<th>L1</th>
</tr>
</thead>
<tbody><tr>
<td>Answer flips when options reversed</td>
<td>0.230</td>
<td><strong>0.073</strong></td>
<td>0.077</td>
</tr>
<tr>
<td>Accuracy</td>
<td>0.747</td>
<td>0.803</td>
<td>0.807</td>
</tr>
<tr>
<td>ECE</td>
<td>0.240</td>
<td>0.184</td>
<td><strong>0.095</strong></td>
</tr>
<tr>
<td><strong>Auto-decidable at ≤5% error</strong></td>
<td><strong>7.7%</strong></td>
<td>46.3%</td>
<td><strong>52.0%</strong></td>
</tr>
</tbody></table>
<p>All of that is on frozen weights, and an independent rerun came back bit-identical on every zero-label number.</p>
<p>The limitations section is exemplary, and it constrains how those numbers should be read: L0 is not a free win on every split, the batch prior needs at least 8 items of the same question, 26 options maximum, L1 does not survive distribution shift, coverage at a 5% risk budget is a high-variance point estimate at n=300, and everything so far is one model family.</p>
<p><strong>Cost:</strong> a K-option choice costs K prefills (2 for <code>noul</code>, 1 for <code>score</code>), all sharing the state prefix and batchable. About 0.25 s per decision at batch 32 on an H100 with 20 permutations. Whether that multiplier is affordable on an M1 Max is an open question.</p>
<p><strong>Runs on:</strong> whatever your base model runs on. No weights of its own.</p>
<h3>12. open-alternative-jev</h3>
<p><a href="https://github.com/ikermoel/open-alternative-jev">github.com/ikermoel/open-alternative-jev</a></p>
<p><em>Read 23 Sep 2026 at</em> <a href="https://github.com/ikermoel/open-alternative-jev/commit/4a85df1831b537343c9e133b5150fa3a3b1ce98e"><code>4a85df1</code></a> <em>· no weights of its own</em></p>
<p>Zero training. The state is written <strong>once</strong> and every question is answered from the next-token distribution at its own packed position.</p>
<p>The packing result is clean (S). On RACE-H, 250 passages × 4 questions: packed scoring gives 92.9% accuracy at 4.55 q/s using 186,898 tokens, against 92.6% at 1.66 q/s and 468,583 tokens for one question per forward pass. That is 2.5× fewer tokens and 2.3× the throughput at the same accuracy.</p>
<p>The headline result is on the 2,000-decision <a href="https://huggingface.co/datasets/LocalLLaMA/typed-decisions"><code>LocalLLaMA/typed-decisions</code></a> suite: <strong>zero-shot Qwen3.6-27B at 73.7% accuracy and ECE 0.020, against Jev's 72.7% and 0.144</strong>. Both sit at the teacher's 73.5% self-agreement ceiling, so the accuracy column is saturated and calibration is the only place left where the models differ. The authors verified their scorer by reproducing laya's published numbers exactly before running anything else, which is what makes the rest credible.</p>
<p><strong>The caution that goes with that number:</strong> JevBench ranked this project using the author's own option order (<code>A. yes, B. no</code>) and found the same model scored <strong>21% instead of 72%</strong> with the options reversed. Read it alongside AnyJev, which exists to remove exactly that failure mode.</p>
<p><strong>Runs on:</strong> your base model. ≤4B bases on a Mac; the 27B result was measured on an H200 MIG slice.</p>
<h3>13. djev-dev</h3>
<p><a href="https://github.com/Davipar/djev-dev">github.com/Davipar/djev-dev</a></p>
<p><em>Read 23 Sep 2026 at</em> <a href="https://github.com/Davipar/djev-dev/commit/3ce907e6835212f27ee82b4cee9039198c4abe35"><code>3ce907e</code></a> <em>· weights</em> <a href="https://huggingface.co/google/diffusiongemma-26B-A4B-it/tree/f7f5b7f5fa82ffc52addd066915886d497f5517b"><code>google/diffusiongemma-26B-A4B-it@f7f5b7f</code></a></p>
<p>Documented, not runnable on a Mac. DiffusionGemma-26B-A4B through vLLM: it denoises all answer positions together rather than scoring a next-token distribution, with native image input and images as options. It adds no weights of its own.</p>
<p>It is the only architecturally different entry in this list, at around #5 on JevBench with 73.0 (T). A 26B MoE plus CUDA-only vLLM puts it out of reach here. If the comparison ever adds a "does the readout architecture matter" arm, this is the entry for it.</p>
<h2>Where these converge, and where they split from Jev</h2>
<p><strong>Same as Jev.</strong> Typed primitives with probabilities read from the model rather than generated as text. One forward pass. Probabilities exposed so you can compute your own confidence. Four of thirteen match the wire format exactly.</p>
<p><strong>Different from Jev, on four axes:</strong></p>
<ol>
<li><p><strong>Trainability.</strong> Jev offers none. Nine of these can be fine-tuned on your own labels.</p>
</li>
<li><p><strong>Calibration as shipped.</strong> kev bakes in T ≈ 2.1–2.4, von ships T = 1.1692, laya ships fitted temperatures for one checkpoint and none for another, nimble ships T = 1.0 untuned, and Jev ships whatever RLCD produced. <strong>Comparing raw ECE across these compares post-processing choices, not models.</strong> I think this is the single biggest reason the existing comparisons do not mean what they appear to mean.</p>
</li>
<li><p><strong>Architecture.</strong> Encoders (laya, von, jeff, GLiNER2), decoder plus LoRA plus head (kev, nimble, decider, NanoJev), pure readout over a stock model (SemIf, AnyJev, open-alternative-jev), and diffusion (djev-dev). Jev is inferred to be a sparse MoE decoder with a linear head.</p>
</li>
<li><p><strong>Scope.</strong> Jev is general text decisions. NanoJev is embodied control. GLiNER2 is extraction. djev-dev is multimodal.</p>
</li>
</ol>
<h3>The shared weakness: option order</h3>
<p><a href="https://archerhume.com/posts/jevs-architecture-unmasked/">Archer Hume</a> measured a −0.28 mean log-odds shift in Jev from adding irrelevant options. JevBench recorded 72% against 21% from reversing two options. AnyJev measured a 0.230 raw flip rate. Nobody is immune, including Jev. Any comparison that runs each choice item in one option order is not measuring what it claims to measure.</p>
<h3>Zero-training readouts are already competitive</h3>
<p>AnyJev and open-alternative-jev, on different bases and different suites, converge on the same finding: a calibrated zero-training readout of a stock open model matches or beats Jev on the probabilities while matching or slightly trailing on accuracy. AnyJev moves auto-decidable traffic 6.8× on frozen weights. open-alternative-jev puts a zero-shot 27B at Jev's accuracy with seven times better ECE. AnyJev's re-measurement of laya on laya's own benchmark has the fine-tuned checkpoint winning on argmax (0.768) while its ECE is six times worse than a calibrated stock model's (0.215 against 0.036).</p>
<p>AnyJev also previews a closed-form head: a shrunk-LDA or ridge matrix fitted on the same hidden state from 200 labels per question, no gradients, seconds on a CPU, reaching 0.771 accuracy and 0.120 ECE against fine-tuned laya's 0.768 and 0.215.</p>
<p>Both results are single-model-family and single-benchmark, and both carry the order-sensitivity caveat above, so I would not treat this as settled. It does raise the question I want the second post to answer: if a properly-read stock model is already at Jev's level, what are the nine trained open decision models for?</p>
<h2>Shortlist A — runs on an M1 Max with 64 GB</h2>
<p>Criteria: open weights or a stock base, runs on MPS or MLX without patching, pinned and licensed, and has a published number I can check. If you want one answer rather than nine, start with <strong>decider-2b</strong>.</p>
<table>
<thead>
<tr>
<th>Model</th>
<th>Download</th>
<th>Published Mac latency</th>
<th>Pick it when</th>
</tr>
</thead>
<tbody><tr>
<td><strong>decider-2b</strong></td>
<td>3.8 GB</td>
<td>133 ms median (M1 Pro)</td>
<td>You want the closest drop-in to Jev's envelope, with Mac numbers someone actually measured rather than asserted</td>
</tr>
<tr>
<td><strong>laya-mlx</strong></td>
<td>2.4 GB</td>
<td>13.4 ms P50 (M3 Max)</td>
<td>Latency dominates and you are willing to fine-tune. Near-chance zero-shot, so this is a base to specialise, not a drop-in</td>
</tr>
<tr>
<td><strong>SemIf</strong> (Qwen3.5-4B)</td>
<td>~8 GB</td>
<td>1.02 s / 21 criteria (3090)</td>
<td>You already run a Qwen 4B and would rather read it properly than add a model</td>
</tr>
<tr>
<td><strong>AnyJev</strong> on a base you have</td>
<td>0</td>
<td>K prefills per K-option choice</td>
<td>You want calibrated, order-stable decisions from existing weights and can afford the prefill multiplier. <strong>26-option cap</strong></td>
</tr>
<tr>
<td><strong>open-alternative-jev</strong> on a base you have</td>
<td>0</td>
<td>packing cuts tokens 2.5×</td>
<td>Many questions against one long state — packing is the cheapest way to ask them</td>
</tr>
<tr>
<td><strong>jeff</strong> (GLiFormer 400M)</td>
<td>~0.8 GB</td>
<td>151 ms p50 (L4)</td>
<td>Cost dominates and you can absorb ~15 accuracy points. No other project here occupies this corner</td>
</tr>
<tr>
<td><strong>kev-0.8b</strong> (or kev-4b@qwen3)</td>
<td>~1.6 GB</td>
<td>329 ms / 174 ms (M5)</td>
<td>You want the harness as much as the model. Use the previous-generation base on a Mac</td>
</tr>
<tr>
<td><strong>von-1.0</strong></td>
<td>3.2 GB</td>
<td>~18 ms claimed</td>
<td>You want a cheap encoder and are prepared to verify its claims yourself</td>
</tr>
<tr>
<td><strong>nimble-9b</strong></td>
<td>~18 GB</td>
<td>444 ms median (M5 Pro 64 GB)</td>
<td>You have 64 GB, ≤26 options and prompts under 2,048 tokens, and you intend to fine-tune</td>
</tr>
</tbody></table>
<p>If you have labels, laya fine-tuned and nimble are the two built to be specialised, and both publish what specialisation bought them. On a smaller Mac, everything except nimble-9b fits in 16 GB.</p>
<p>One practical warning about dependencies: <code>decider-ai</code> pins <code>numpy&lt;2</code> and <code>transformers&gt;=5</code>, laya wants <code>transformers&gt;=4.48</code>, and laya-mlx wants neither. <strong>One virtualenv per model is mandatory.</strong> nimble's README already enforces three for the same reason.</p>
<p>These recommendations rest entirely on other people's numbers, and the second post is where I find out whether I would still give the same answer afterwards.</p>
<h2>Shortlist B — needs a real GPU</h2>
<table>
<thead>
<tr>
<th>Model</th>
<th>Why it will not fit</th>
<th>What it would answer</th>
</tr>
</thead>
<tbody><tr>
<td><strong>decider-35b-a3b</strong></td>
<td>69 GB bf16, exceeds 64 GB unified; the NVFP4 build is Blackwell-targeted</td>
<td>Does scale close the knowledge gap for a trained model?</td>
</tr>
<tr>
<td><strong>open-alternative-jev on Qwen3.6-27B</strong></td>
<td>8-bit needs ~35 GB; measured on an H200 MIG slice</td>
<td>Does scale close it <em>without</em> training?</td>
</tr>
<tr>
<td><strong>djev-dev</strong></td>
<td>26B MoE plus CUDA-only vLLM</td>
<td>Does the readout architecture itself matter?</td>
</tr>
<tr>
<td><strong>kev-9b / nimble-9b for latency</strong></td>
<td>They run on a Mac for correctness, but Mac latency numbers are not GPU latency numbers</td>
<td>Fair latency on known hardware</td>
</tr>
<tr>
<td><strong>NanoJev</strong></td>
<td>CUDA serve plus ViZDoom, and it is a games model</td>
<td>Nothing this comparison is asking</td>
</tr>
</tbody></table>
<p>None of these is required for the core comparison. One thing to know before you plan around free compute: two T4s do not pool into 32 GB, so Kaggle is good for ≤4B models, encoders and latency measurement on a known GPU, and is not a general escape hatch.</p>
<h2>Key performance results</h2>
<ul>
<li><p><a href="https://github.com/ikermoel/open-alternative-jev"><strong>open-alternative-jev</strong></a> — Slightly better accuracy than Jev: <strong>73.7% vs 72.7%</strong>, with better-calibrated probabilities.</p>
</li>
<li><p><a href="https://github.com/nokia-applied-research/AnyJev"><strong>AnyJev</strong></a> — Achieves <strong>similar accuracy to Jev without fine-tuning</strong>, and also produces well-calibrated confidence scores.</p>
</li>
<li><p><a href="https://github.com/Davipar/djev-dev"><strong>djev-dev</strong></a> — Performs very close to Jev on JevBench: <strong>~73 vs ~74–75</strong>.</p>
</li>
<li><p><a href="https://github.com/bespokelabsai/nimble"><strong>Nimble</strong></a> — A smaller model that still gets close to Jev: <strong>90.1% vs 93.2%</strong> on its benchmark.</p>
</li>
<li><p><a href="https://github.com/TheoLeeCJ/SemIf-OpenJev"><strong>SemIf-OpenJev</strong></a> — Uses <strong>no fine-tuning</strong> and still ranks around <strong>#2–#4 on JevBench</strong>.</p>
</li>
</ul>
<h2>Open questions, and what the second post will do</h2>
<p><strong>Assumptions in this post.</strong> Every number is published by someone else, and star counts and rankings were read on 22–23 September 2026. JevBench's Jev score already differs between my two reads, 75.4 and 74.4, which tells you how fast this is moving.</p>
<p><strong>What still looks open, as far as I can find:</strong></p>
<ul>
<li><p>I have not found a published run that puts the trained open models through one equal-calibration protocol on one frozen suite. If someone has done it since, that is the result I most want pointed at.</p>
</li>
<li><p>I have not seen several of these run on <strong>one</strong> Mac with <strong>one</strong> request shape. The published Mac numbers I could find span four different Macs: M1 Pro, M3 Max, M5 and M5 Pro.</p>
</li>
<li><p>"Does it know it doesn't know" is the least crowded question. Jev answers at ≥0.9 confidence on 9% of items whose deciding evidence was removed, kev-9b on 5%, kev-8b on 26%. <a href="https://github.com/scienthoon/jev-ood-calibration">scienthoon's unknowable-priority suite</a> pins Jev at 44.7% accuracy while it states a mean probability of 0.74 on the level it chose.</p>
</li>
<li><p>How much of any gap is the readout rather than the model.</p>
</li>
</ul>
<p><strong>The second post:</strong> one machine, one frozen suite, the same calibration budget for every model, at least two option orders per choice item, latency reported in families that are never combined into one ranking, and raw responses published.</p>
<p>Roughly forty projects appeared in a week, and any ranking published today is stale in two, which is why the second post is a protocol rather than a scoreboard.</p>
<p>If I have mis-read a project or graded a number wrongly, I would like to know. Several of these authors document their own weaknesses better than most commercial vendors do, and they deserve to be quoted accurately.</p>
<h2>Appendix A — sources</h2>
<p>Every number above traces to one of these. All public, all read on 22–23 September 2026.</p>
<p><strong>TypeSafe (class V)</strong></p>
<ul>
<li><p><a href="https://typesafe.ai/">typesafe.ai</a></p>
</li>
<li><p><a href="https://docs.typesafe.ai/introduction/quickstart">docs.typesafe.ai — quickstart</a></p>
</li>
<li><p><a href="https://docs.typesafe.ai/model-jaggedness/jev-1.13">docs.typesafe.ai — model jaggedness, jev-1.13</a></p>
</li>
</ul>
<p><strong>Independent analyses of Jev (class T)</strong></p>
<ul>
<li><p><a href="https://archerhume.com/posts/jevs-architecture-unmasked/">Archer Hume — Jev's architecture unmasked</a> — black-box probing, tokenizer matching against 192 candidates, option-interaction experiments, MMLU ECE 0.0313</p>
</li>
<li><p><a href="https://amaarora.github.io/posts/2026-19-09-jev-intro.html">Aman Arora — a first look at Jev</a> — primitives, context split, 299 ms against 1,114 ms</p>
</li>
<li><p><a href="https://amaarora.github.io/posts/2026-21-09-oss-jev-interface.html">Aman Arora — an open-source Jev interface</a></p>
</li>
<li><p><a href="https://arcturus-labs.com/blog/2026/09/16/typesafes-jev-trades-text-generation-for-instant-calibrated-decisions/">Arcturus Labs — Jev trades text generation for instant calibrated decisions</a> — coin-flip calibration probe</p>
</li>
<li><p><a href="https://anth.us/blog/can-you-trust-jev-confidence/">anth.us — can you trust Jev confidence?</a> — AUROC 0.83, isotonic ECE 0.008</p>
</li>
</ul>
<p><strong>The projects</strong></p>
<p><a href="https://github.com/NandhaKishorM/laya">laya</a> · <a href="https://github.com/mizorewww/laya-mlx">laya-mlx</a> · <a href="https://github.com/jaredpalmer/kev">kev</a> · <a href="https://github.com/TheoLeeCJ/SemIf-OpenJev">SemIf</a> · <a href="https://github.com/fastino-ai/GLiNER2">GLiNER2</a> · <a href="https://github.com/TianyuCodings/NanoJev">NanoJev</a> · <a href="https://github.com/bespokelabsai/nimble">nimble</a> · <a href="https://github.com/wfzyx/von">von</a> · <a href="https://github.com/Mapika/decider">decider</a> · <a href="https://github.com/logan-markewich/jeff">jeff</a> · <a href="https://github.com/nokia-applied-research/AnyJev">AnyJev</a> · <a href="https://github.com/ikermoel/open-alternative-jev">open-alternative-jev</a> · <a href="https://github.com/Davipar/djev-dev">djev-dev</a></p>
<p><strong>Benchmarks, harnesses and datasets</strong></p>
<ul>
<li><p><a href="https://github.com/fstandhartinger/jevbench">JevBench</a> — 534 frozen decisions, 66 ranked rows, hard tier hashed before any run</p>
</li>
<li><p><a href="https://github.com/apolinario/decision-index">Decision Index</a> — 132,422 requests over 37 benchmarks, <code>transformers</code> engine with <code>device=mps</code></p>
</li>
<li><p><a href="https://github.com/scienthoon/jev-ood-calibration">jev-ood-calibration</a> — a regenerable 900-ticket synthetic suite with an unknowable-by-design label</p>
</li>
<li><p><a href="https://github.com/AbdelStark/jev-benchmarks">jev-benchmarks</a> · <a href="https://github.com/nibzard/decision-model-benchmark">decision-model-benchmark</a> · <a href="https://github.com/TianyuCodings/JevHarness">JevHarness</a></p>
</li>
<li><p><a href="https://huggingface.co/datasets/LocalLLaMA/typed-decisions"><code>LocalLLaMA/typed-decisions</code></a> — 400 cases × 5 typed questions, the de-facto community suite</p>
</li>
<li><p><a href="https://huggingface.co/datasets/Luni/laya-jev-benchmark"><code>Luni/laya-jev-benchmark</code></a> — third-party scorer with a Jev 1.13.0 row</p>
</li>
</ul>
<p><strong>Discovery aids</strong></p>
<ul>
<li><p><a href="https://systemonemodels.org/examples/alternatives/">systemonemodels.org/examples/alternatives</a> — 20 listed reproductions, but it misreports kev, decider and SemIf as having no published metrics. Use it to find things, not to cite them.</p>
</li>
<li><p>GitHub <code>topic:system-one</code> and <code>topic:jev</code>.</p>
</li>
</ul>
<p><strong>Unverified, and flagged as such</strong></p>
<ul>
<li><p>Reports of "$5 free credit, no waitlist" for the TypeSafe console are secondary and SEO-flavoured; typesafe.ai still renders a waitlist element.</p>
</li>
<li><p>nimble's repository has no LICENSE file although its model card says Apache-2.0. NanoJev's Hugging Face model-card licence is unset.</p>
</li>
<li><p>JevBench's Jev score differs between my two reads (75.4 and 74.4). Re-read before citing.</p>
</li>
</ul>
<h2>Appendix B — live additions</h2>
<p>List of useful Jev related links as discovered.</p>
<table>
<thead>
<tr>
<th>Added</th>
<th>Detail</th>
</tr>
</thead>
<tbody><tr>
<td>25 Sep</td>
<td>Jev replications, models, interpretations and papers in one place:</td>
</tr>
<tr>
<td><a href="https://hanxiao.io/all-about-jev/#v=table&amp;cat=model&amp;sort=score&amp;d=-1">https://hanxiao.io/all-about-jev/#v=table&amp;cat=model&amp;sort=score&amp;d=-1</a></td>
<td></td>
</tr>
<tr>
<td></td>
<td></td>
</tr>
</tbody></table>
]]></content:encoded></item><item><title><![CDATA[Lecture summary: Trends in AI by Jeff Dean]]></title><description><![CDATA[Jeff Dean, Chief Scientist for Google Research and Google DeepMind presents his talk "Important Trends in AI: How Did We Get Here, What Can We Do Now, and Where are We Headed?" as part of the Princeto]]></description><link>https://infocruncher.hashnode.dev/lecture-summary-trends-in-ai-by-jeff-dean</link><guid isPermaLink="true">https://infocruncher.hashnode.dev/lecture-summary-trends-in-ai-by-jeff-dean</guid><category><![CDATA[lecture-summary]]></category><dc:creator><![CDATA[Dylan Hogg]]></dc:creator><pubDate>Wed, 23 Sep 2026 01:42:46 GMT</pubDate><content:encoded><![CDATA[<p>Jeff Dean, Chief Scientist for Google Research and Google DeepMind presents his talk "<a href="https://www.youtube.com/watch?v=UTTeXZrpMR0">Important Trends in AI: How Did We Get Here, What Can We Do Now, and Where are We Headed?</a>" as part of the Princeton CS Distinguished Colloquia series on February 10, 2026.</p>
<h2>Summary</h2>
<p>This is a broad technical/history lecture by <strong>Jeff Dean</strong>, described in the introduction as Google’s chief scientist and a co-lead of the Gemini project. His central argument is that modern AI did <strong>not emerge from one breakthrough</strong>. It is the cumulative result of improvements in <strong>scale, algorithms, architectures, hardware, distributed systems, training methods and inference techniques</strong>, with these improvements multiplying rather than merely adding together.</p>
<p>The talk moves from Dean's early neural-network work through distributed training, word embeddings, seq2seq, TPUs, Transformers, sparse/MoE models, Pathways, inference-time reasoning, distillation, reinforcement learning and speculative decoding, before explaining how these pieces come together in Gemini and where he thinks AI research is heading.</p>
<p>The talk in one sentence: Modern AI is best understood as a co-evolving stack—models, algorithms, data, distributed systems and specialised hardware—and the next phase shifts much of the emphasis from merely training bigger models toward efficient inference, long-term retrieval/memory, multimodal world models, reasoning, and coordinated populations of AI agents.</p>
<h2>Key points</h2>
<ol>
<li><p><strong>Scale remains important, but scale alone is the wrong explanation for recent AI progress.</strong> Dean argues that increasing compute/data/model size has produced fairly continuous gains for roughly 13–14 years, but architectural and algorithmic improvements have been equally important. His example is that a 20× scaling improvement combined with a 50× algorithmic improvement can yield something closer to a 1,000× overall improvement.</p>
</li>
<li><p><strong>Distributed training was foundational very early.</strong> Dean's undergraduate work already explored what would now be called model and data parallelism. At Google, this developed into systems capable of training neural networks <strong>50–100× larger</strong> than contemporary models using many asynchronous model replicas and distributed parameter servers.</p>
</li>
<li><p><strong>Representation learning showed that models could discover meaningful concepts without explicit labels.</strong> Google's large unsupervised vision experiments on 10 million YouTube frames produced high-level units responsive to things such as cats, faces and people. Using the learned representation to initialise supervised training produced a large improvement on ImageNet-22K. Word-vector work similarly demonstrated that semantic relationships naturally emerge geometrically in embedding space.</p>
</li>
<li><p><strong>Specialised AI hardware became necessary because inference itself was becoming economically impossible on conventional compute.</strong> Dean describes estimating that deploying a much better speech model to one billion users for three minutes per day would require roughly <strong>doubling Google's entire computer fleet</strong>. That helped motivate TPUs. TPUv1 exploited reduced precision and the fact that neural networks are dominated by a relatively small set of linear-algebra operations.</p>
<p>In the figures given in the talk, TPUv1 was <strong>15–30× faster</strong> and <strong>30–80× more energy efficient</strong> than contemporary CPUs/GPUs. Later TPU systems evolved into interconnected ML supercomputers; Dean says Ironwood delivers roughly <strong>3,600× the pod-level performance of TPUv2</strong> and ~30× the FLOPs/watt.</p>
</li>
<li><p><strong>Transformers solved two major limitations of recurrent language models.</strong> LSTMs were inherently sequential and compressed everything encountered so far into one state vector. Transformers instead retain representations and use learnable attention to access them, making computation much more parallelisable and information easier to retrieve. Dean presents the Transformer as dramatically more compute-efficient than comparable LSTMs.</p>
</li>
<li><p><strong>Self-supervised learning unlocked essentially unlimited training data.</strong> Predicting the next token or masked tokens gives a precise training signal without requiring humans to label data. Autoregressive prediction naturally became the basis for generative/chat models, while fill-in-the-blank objectives remain particularly useful for representation learning.</p>
</li>
<li><p><strong>Sparse / mixture-of-experts models are a major part of the scaling story.</strong> Rather than activating every parameter for every token, sparse models learn specialised components and a routing mechanism that selects which experts to activate. This allows much greater model capacity without proportional inference cost. Dean cites an example giving around an <strong>8× reduction in training compute for equivalent accuracy</strong>.</p>
</li>
<li><p><strong>At extreme scale, reliability becomes an ML problem as well as a systems problem.</strong> Google's Pathways attempts to present thousands of accelerators as one giant computer. At this scale, silent data corruption becomes significant: hardware can occasionally return incorrect values rather than simply fail. Google monitors training signals such as gradient norms, uses deterministic replay to distinguish anomalous data from hardware errors, and swaps faulty hardware for hot spares.</p>
</li>
<li><p><strong>Inference-time compute is now another scaling axis.</strong> Dean frames chain-of-thought-style reasoning as giving the model additional computation at inference time: each additional generated token represents another model pass. This substantially improved mathematical problem solving once models became sufficiently capable.</p>
<p>The later Q&amp;A extends this idea: for critical tasks, generate multiple candidate solutions, have the model evaluate them, and retain the strongest result. Dean says this can reduce hallucination rates, but it makes <strong>inference efficiency even more important</strong>.</p>
</li>
<li><p><strong>Distillation is central to making powerful models economical.</strong> Instead of training a smaller model only against hard labels, the teacher's full probability distribution supplies much richer supervision. Dean explicitly says this is used in Gemini to move capability from <strong>Pro-scale models into smaller Flash-scale models</strong>.</p>
</li>
<li><p><strong>RL/post-training increasingly determines what capabilities a pretrained model actually exhibits.</strong> Reward signals can come from humans, another model, or objectively verifiable outcomes such as compiling code, passing unit tests or proving a theorem. Dean identifies <strong>RL for non-verifiable domains</strong> as an important unresolved research problem: where do sufficiently reliable reward signals come from?</p>
</li>
<li><p><strong>Speculative decoding addresses the sequential bottleneck of autoregressive generation.</strong> A cheap draft model proposes several tokens, while the large target model checks them in parallel. Importantly, Dean stresses that this requires no architecture change or retraining and preserves the target model's output distribution while improving utilisation.</p>
</li>
</ol>
<h2>Gemini: the synthesis of these ideas</h2>
<p>Dean describes Gemini as an attempt beginning in <strong>February 2023</strong> to consolidate previously separate Google language and multimodal efforts into one major project. The design goal was <strong>multimodality from the beginning</strong>: text, images, audio and video, with newer systems also exposed to things such as LiDAR and robotic-control data.</p>
<p>Gemini therefore isn't presented as one novel architecture. It is a system incorporating essentially the whole preceding history:</p>
<p><strong>TPUs + distributed/model/data parallelism + cross-datacenter training + Pathways + JAX + Transformers + sparse models + distillation + long context + inference-time reasoning + speculative decoding + SFT + RL.</strong></p>
<p>One particularly interesting point is Dean's view of <strong>context vs model weights</strong>. Information absorbed into parameters from trillions of training tokens becomes somewhat "muddled and fuzzy"; information supplied directly in context remains precise. That is part of Google's motivation for pushing very long context windows.</p>
<p>Google's desired model progression is also notable:</p>
<blockquote>
<p><strong>Next-generation Flash ≳ previous-generation Pro</strong></p>
</blockquote>
<p>The idea is that frontier-level capability should repeatedly migrate into a much cheaper model tier, making yesterday's expensive capabilities economical for mainstream applications.</p>
<h2>Capabilities he highlights</h2>
<p>The talk uses several examples to illustrate how rapidly the nature of model capability is changing:</p>
<ul>
<li><p><strong>Mathematical reasoning:</strong> a general-purpose Gemini Pro-scale model, given a large inference-time thinking budget, achieved a gold-medal-level result on the IMO rather than relying on the previous collection of specialised theorem-proving/geometry systems.</p>
</li>
<li><p><strong>Generative UI/software:</strong> a model can take source material and dynamically construct an interactive visual explanation or complete website. Dean expects this to mean <strong>far more software gets created</strong>, including by people unable to program conventionally.</p>
</li>
<li><p><strong>Visual reasoning:</strong> image-generation models can reason in "pixel space", generating intermediate visual states of a physical problem rather than reasoning solely through language.</p>
</li>
<li><p><strong>World models:</strong> he treats world modelling as closely related to multimodal understanding. Gemini-derived systems can generate persistent interactive worlds, which can in turn produce unusual simulated scenarios for systems such as autonomous vehicles.</p>
</li>
</ul>
<h1>Most important forward-looking ideas</h1>
<h3>1. Humans managing teams of AI agents</h3>
<p>Dean expects the dominant interaction model to move beyond:</p>
<p><strong>1 person → 1 chatbot</strong></p>
<p>towards something more like:</p>
<p><strong>1 person → dozens or hundreds of AI agents</strong></p>
<p>That raises new problems around HCI, delegation, coordination, communication between agents and converting relatively weakly specified human goals into reliable execution.</p>
<p>This is probably the strongest forward-looking theme in the talk.</p>
<h3>2. Context windows alone won't solve long-term memory</h3>
<p>A million tokens is useful, but Dean asks what happens when the useful information is effectively <strong>a trillion tokens</strong>.</p>
<p>His likely architecture is hybrid:</p>
<p><strong>huge corpus → learned retrieval → lightweight relevance filtering → small highly relevant subset → context window → model</strong></p>
<p>Potential applications include a Gemini able—with permission—to reason over all of someone's email/photos, multimodal search across huge video collections, or coding agents able to draw upon an entire corporate codebase.</p>
<p>This is effectively a prediction that <strong>retrieval and context management become core parts of the model system</strong>, rather than endlessly increasing raw context length.</p>
<h3>3. Inference efficiency may matter more than training efficiency</h3>
<p>As agents proliferate and invoke other agents, inference volume compounds. Dean therefore expects:</p>
<ul>
<li><p>specialised inference hardware,</p>
</li>
<li><p>better inference algorithms,</p>
</li>
<li><p>aggressive latency optimisation,</p>
</li>
<li><p>AI-assisted chip design.</p>
</li>
</ul>
<p>He specifically contrasts an interaction responding in roughly <strong>100 ms versus 5 seconds</strong>: latency becomes a product capability, not merely an infrastructure metric.</p>
<h3>4. AI research does not necessarily require frontier-scale compute</h3>
<p>His final answer is particularly relevant for researchers without thousands of accelerators. He recommends:</p>
<ul>
<li><p>test genuinely different ideas at <strong>very small scale</strong>;</p>
</li>
<li><p>run several scales;</p>
</li>
<li><p>study the <strong>slope/scaling trend</strong>, rather than obsessing over absolute benchmark position;</p>
</li>
<li><p>favour novel approaches with promising scaling behaviour over tiny improvements to the current SOTA.</p>
</li>
</ul>
<p>An idea slightly below the baseline at tiny scale but improving faster may be much more important than an idea that narrowly beats the baseline at one small scale.</p>
<p>That is essentially an argument for <strong>researching scaling behaviour rather than leaderboard points</strong>.</p>
<h2>His overall view of AI risk and impact</h2>
<p>Dean is relatively optimistic about catastrophic AI safety concerns, saying he believes careful engineering can constrain what systems are permitted to do. His nearer-term concerns are more concrete:</p>
<ul>
<li><p>highly convincing AI-generated misinformation/audio/video;</p>
</li>
<li><p>managing labour and skill transitions;</p>
</li>
<li><p>ensuring people learn to use AI tools rather than simply being displaced by automation.</p>
</li>
</ul>
<p>He expects major effects across <strong>employment, education, healthcare, misinformation/media, governance/national security, entertainment and AI-for-science</strong>, and sees widespread access to previously scarce expertise as one of the major benefits.</p>
<h2>Appendix: Further reading</h2>
<p><strong>Overview</strong></p>
<ul>
<li><a href="https://www.amacad.org/publication/daedalus/golden-decade-deep-learning-computing-systems-applications"><strong>A Golden Decade of Deep Learning: Computing Systems &amp; Applications</strong></a> (Jeff Dean, <em>Daedalus</em>, 2022) — Dean's own written account of much of the same history: scale, hardware, and the applications they unlocked.</li>
</ul>
<p><strong>Early foundations: distributed training and representation learning</strong></p>
<ul>
<li><p><a href="https://papers.nips.cc/paper/2012/hash/6aca97005c68f1206823815f66102863-Abstract.html"><strong>Large Scale Distributed Deep Networks</strong></a> (Dean et al., NeurIPS 2012) — the DistBelief paper: asynchronous model replicas and parameter servers for training much larger networks.</p>
</li>
<li><p><a href="https://arxiv.org/abs/1112.6209"><strong>Building High-level Features Using Large Scale Unsupervised Learning</strong></a> (Le et al., 2012) — the YouTube-frames "cat neuron" experiment and its ImageNet-22K result.</p>
</li>
<li><p><a href="https://arxiv.org/abs/1301.3781"><strong>Efficient Estimation of Word Representations in Vector Space</strong></a> (Mikolov et al., 2013) — word2vec, where semantic relationships emerge geometrically in embedding space.</p>
</li>
<li><p><a href="https://arxiv.org/abs/1409.3215"><strong>Sequence to Sequence Learning with Neural Networks</strong></a> (Sutskever, Vinyals &amp; Le, 2014) — the LSTM encoder-decoder approach that Transformers later superseded.</p>
</li>
</ul>
<p><strong>Hardware and systems</strong></p>
<ul>
<li><p><a href="https://arxiv.org/abs/1704.04760"><strong>In-Datacenter Performance Analysis of a Tensor Processing Unit</strong></a> (Jouppi et al., 2017) — the TPUv1 paper, source of the speed and energy-efficiency comparisons against CPUs/GPUs.</p>
</li>
<li><p><a href="https://blog.google/products/google-cloud/ironwood-tpu-age-of-inference/"><strong>Ironwood: The first Google TPU for the age of inference</strong></a> (Google, 2025) — the latest TPU generation Dean cites, and its inference-first framing.</p>
</li>
<li><p><a href="https://arxiv.org/abs/2203.12533"><strong>Pathways: Asynchronous Distributed Dataflow for ML</strong></a> (Barham et al., 2022) — the system for presenting thousands of accelerators as one computer.</p>
</li>
</ul>
<p><strong>Architectures and efficiency techniques</strong></p>
<ul>
<li><p><a href="https://arxiv.org/abs/1706.03762"><strong>Attention Is All You Need</strong></a> (Vaswani et al., 2017) — the original Transformer paper, replacing recurrence with attention.</p>
</li>
<li><p><a href="https://arxiv.org/abs/1701.06538"><strong>Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer</strong></a> (Shazeer et al., 2017) — the sparse MoE layer with learned routing that underpins the sparse-model scaling story.</p>
</li>
<li><p><a href="https://arxiv.org/abs/1503.02531"><strong>Distilling the Knowledge in a Neural Network</strong></a> (Hinton, Vinyals &amp; Dean, 2015) — the distillation technique Dean describes for moving Pro-scale capability into Flash-scale models.</p>
</li>
<li><p><a href="https://arxiv.org/abs/2211.17192"><strong>Fast Inference from Transformers via Speculative Decoding</strong></a> (Leviathan, Kalman &amp; Matias, 2022) — draft-and-verify decoding that speeds up generation without changing the output distribution.</p>
</li>
</ul>
<p><strong>Reasoning and Gemini</strong></p>
<ul>
<li><p><a href="https://arxiv.org/abs/2201.11903"><strong>Chain-of-Thought Prompting Elicits Reasoning in Large Language Models</strong></a> (Wei et al., 2022) — the paper behind treating intermediate reasoning tokens as extra inference-time compute.</p>
</li>
<li><p><a href="https://arxiv.org/abs/2312.11805"><strong>Gemini: A Family of Highly Capable Multimodal Models</strong></a> (Gemini Team, 2023) — the technical report for the natively multimodal model that pulls these ideas together.</p>
</li>
<li><p><a href="https://deepmind.google/discover/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/"><strong>Gemini Deep Think achieves gold-medal standard at the IMO</strong></a> (Google DeepMind, 2025) — the IMO result Dean highlights, reached by a general model with a large thinking budget.</p>
</li>
</ul>
]]></content:encoded></item><item><title><![CDATA[Embedding model dimensions for image search: what the published evidence shows]]></title><description><![CDATA[Publicly published results suggest that embedding vectors often add cost without meaningfully improving image-search quality, so smaller dimensions are worth evaluating.
1. Two common assumptions abou]]></description><link>https://infocruncher.hashnode.dev/embedding-model-dimensions-for-image-search-what-the-published-evidence-shows</link><guid isPermaLink="true">https://infocruncher.hashnode.dev/embedding-model-dimensions-for-image-search-what-the-published-evidence-shows</guid><category><![CDATA[ai search]]></category><category><![CDATA[Embedding Models]]></category><dc:creator><![CDATA[Dylan Hogg]]></dc:creator><pubDate>Wed, 23 Sep 2026 01:12:29 GMT</pubDate><content:encoded><![CDATA[<p>Publicly published results suggest that embedding vectors often add cost without meaningfully improving image-search quality, so smaller dimensions are worth evaluating.</p>
<h2>1. Two common assumptions about embedding size</h2>
<p>Many of us bring two assumptions to choosing an embedding size.</p>
<p>The first assumption: more dimensions means a richer vector. Twice the coordinates should mean twice the room for visual detail, so 3072 dimensions must hold more of a photograph than 768 does.</p>
<p>The second is the cross-model version: Gemini Embedding 2 outputs 3072 dimensions and some other model outputs 768, so Gemini has more capacity to work with.</p>
<p>The first assumption has a measurable cost. Take the example I will return to throughout this post: a real-estate search system over 10 million property photos. At float32, 3072-d vectors take 122.88 GB of raw payload before any index overhead. The same catalogue at 768-d takes 30.72 GB. That is roughly a 92 GB difference from a single architectural choice you probably made in an afternoon.</p>
<p>If the assumption holds, that extra storage buys better search results. If it does not, the storage buys nothing.</p>
<h2>2. The dimension/quality questions</h2>
<blockquote>
<p>When the <em>same model</em> is served at 3072, 1024 and 256 dimensions, what happens to text→image retrieval quality?</p>
<p>And does the published evidence cover the models you are most likely to deploy?</p>
</blockquote>
<p>The first question has a good answer, from two independent labs. The second does not: for <strong>Gemini Embedding 2 and Cohere Embed 4 — likely two models on your shortlist — nobody has published the per-dimension curve for image retrieval.</strong> I could not find one at all. That gap is half of what this post is about, so I want to state it up front rather than save it for the end.</p>
<h2>3. The one commercial ablation that answers the question: Amazon Nova MME</h2>
<p>Amazon Nova Multimodal Embeddings is the only major commercial multimodal model I found with a published per-dimension table for text→image retrieval. The architecture, training and benchmarks are held constant; only the vector length changes. Amazon states that the shorter vectors are prefixes of the full 3072-d representation, trained so the front of the vector carries the most signal. Section 5 explains why that works.</p>
<table>
<thead>
<tr>
<th>Nova MME dims</th>
<th>TextCaps</th>
<th>MSCOCO</th>
<th>ViDoRe v2</th>
<th>float32/vector</th>
</tr>
</thead>
<tbody><tr>
<td>3072</td>
<td>88.9</td>
<td>76.7</td>
<td>58.7</td>
<td>12 KiB</td>
</tr>
<tr>
<td>1024</td>
<td>87.9</td>
<td>75.6</td>
<td>57.7</td>
<td>4 KiB</td>
</tr>
<tr>
<td>384</td>
<td>85.6</td>
<td>72.9</td>
<td>53.4</td>
<td>1.5 KiB</td>
</tr>
<tr>
<td>256</td>
<td>83.1</td>
<td>70.6</td>
<td>50.2</td>
<td>1 KiB</td>
</tr>
</tbody></table>
<p>Three results stand out.</p>
<p><strong>3072 → 1024 is cheap.</strong> Removing two-thirds of the coordinates costs 1.0 points on TextCaps and 1.1 on MSCOCO.</p>
<p><strong>1024 → 384 costs noticeably more.</strong> Against the 3072 baseline: 3.3 points on TextCaps, 3.8 on MSCOCO, 5.3 on ViDoRe.</p>
<p><strong>256 loses substantial quality.</strong> It costs 5.8, 6.1 and 8.5 points respectively, or roughly 6.5%, 8.0% and 14.5% in relative terms.</p>
<p>The three benchmarks do not degrade at the same rate. ViDoRe (Visual Document Retrieval, a benchmark of document screenshots with dense text and fine visual detail) loses 8.5 points at 256 dimensions where TextCaps loses 5.8. This is the first sign that visually rich retrieval needs more dimensions than object-centric captioning benchmarks reveal. I think that matters for property photos, and section 7 returns to it.</p>
<p>One note on metrics if you plan to reproduce any of this: the image-retrieval columns are average Recall@1/5/10 (R@K is the share of queries where the correct image appears in the top K), and ViDoRe is NDCG@5, a graded ranking score over the top five results. Do not compare these figures against benchmarks that quote Recall@1 alone.</p>
<h2>4. Open-weight corroboration: Jina CLIP v2</h2>
<p>One vendor's ablation of its own model is not a pattern. Jina CLIP v2 serves as a control: the weights are open, text→image Recall@5 is published at six lengths, and you can rerun the evaluation yourself.</p>
<table>
<thead>
<tr>
<th>Dims</th>
<th>CLIP Benchmark</th>
<th>Crossmodal-3600</th>
<th>XTD10</th>
</tr>
</thead>
<tbody><tr>
<td>1024</td>
<td>79.10</td>
<td>81.43</td>
<td>84.87</td>
</tr>
<tr>
<td>768</td>
<td>79.12</td>
<td>82.35</td>
<td>84.85</td>
</tr>
<tr>
<td>512</td>
<td>78.93</td>
<td>82.31</td>
<td>84.60</td>
</tr>
<tr>
<td>256</td>
<td>78.32</td>
<td>81.75</td>
<td>84.32</td>
</tr>
<tr>
<td>128</td>
<td>75.90</td>
<td>78.17</td>
<td>81.80</td>
</tr>
<tr>
<td>64</td>
<td>70.51</td>
<td>72.52</td>
<td>77.85</td>
</tr>
</tbody></table>
<p><strong>768 dimensions is not worse than 1024 here, and on Crossmodal-3600 it is nearly a point better.</strong> Even 256 stays within about one point of full width on all three tests. Quality drops sharply at 128 (3-4 points below full width) and 64 (7-9 points).</p>
<p>This is the same curve shape as Nova's, at a different scale, from a different lab, on different benchmarks. Two independent sources agreeing is the main argument of this post.</p>
<h2>5. Why the curve has this shape</h2>
<p>Both models are trained with <strong>Matryoshka Representation Learning (MRL)</strong>. The training objective is applied not only to the full vector but to nested prefixes of it, so the first 256 coordinates are optimised to work as an embedding on their own, the first 768 likewise, and so on. This means that <strong>the dimensions of an MRL-trained model are not equally informative independent coordinates.</strong> Truncating is not compressing a flat 3072-d space; it is taking a prefix of a representation designed to survive truncation.</p>
<p>This also explains why the cross-model assumption in section 1 fails. Gemini Embedding 2 at 768 dimensions is still the full Gemini network doing the work, which is not the same as swapping in a smaller 768-d vision-language architecture. <strong>Dimension count is not comparable across model families</strong>, only within one.</p>
<p>One reproducibility trap: taking a prefix changes the vector's norm. If you compare an unnormalised truncation against a normalised full vector, you will manufacture a dimensionality effect that is not there. Gemini's API re-normalises automatically when you request a non-default size. If you truncate by hand, L2-normalise before any cosine comparison.</p>
<h2>6. The gap: the two models you may be evaluating</h2>
<p>The missing evidence below is, I think, the most useful thing this post has to say.</p>
<table>
<thead>
<tr>
<th>Model</th>
<th>Dimension choices</th>
<th>Published multimodal per-dimension ablation?</th>
</tr>
</thead>
<tbody><tr>
<td>Gemini Embedding 2</td>
<td>128-3072; recommended 768 / 1536 / 3072</td>
<td><strong>No</strong></td>
</tr>
<tr>
<td>Cohere Embed 4</td>
<td>256 / 512 / 1024 / 1536</td>
<td><strong>No</strong></td>
</tr>
<tr>
<td>Amazon Nova MME</td>
<td>256 / 384 / 1024 / 3072</td>
<td><strong>Yes</strong></td>
</tr>
<tr>
<td>Jina CLIP v2</td>
<td>64 / 128 / 256 / 512 / 768 / 1024</td>
<td><strong>Yes</strong></td>
</tr>
</tbody></table>
<p><strong>Gemini Embedding 2.</strong> What <em>is</em> published is genuinely good: the model is natively multimodal rather than a text model bolted onto a vision encoder, MRL losses are trained at 768- and 1536-d prefixes, and absolute text→image Recall@1 is strong (mean 80.5; DOCCI 93.4, TextCaps 89.6, MSCOCO 62.9). DOCCI (Descriptions of Connected and Contrasting Images) and TextCaps are detailed-caption datasets rather than short object labels. I expect detailed-caption retrieval to map better onto a query like <em>"sunlit open-plan kitchen with a waterfall-edge stone island and black pendant lights"</em> than COCO's dominant-object framing does, but that is my inference about the application rather than something Google measured.</p>
<p>Google has <strong>not</strong> published any 3072 vs 1536 vs 768 text→image table. The dimension ablation in Google's docs covers <strong>Gemini Embedding 001, a text model, evaluated on MTEB</strong> (1536: 68.17, 768: 67.99, 256: 66.19, 128: 63.31). That is encouraging for MRL in general and consistent with everything above, but it is not evidence about Embedding 2 on complex images, and I would rather say so than treat it as a substitute.</p>
<p><strong>Cohere Embed 4.</strong> Embed 4 is genuinely multimodal, with clean MRL sizes. Its distinctive capability is fusing image and text into a single vector, which is attractive for listings where room type, suburb or other structured attributes are reliable text that you would otherwise force the pixels to carry. There is again no published per-dimension image-retrieval curve.</p>
<p>One external datapoint exists, and it needs careful handling. Amazon's Nova report benchmarked Embed 4 through the public Bedrock API and scored it <strong>22.9 on MSCOCO against Nova's 76.7</strong>. Three caveats come with that number: Amazon is a direct competitor, the figures are not Cohere self-reported, and the Embed 4 output dimension used is not stated. I read it as a reason to benchmark Cohere yourself, not as evidence about Cohere's dimension curve, which it says nothing about.</p>
<h2>7. Input resolution probably matters more than dimension</h2>
<p>I have spent a thousand words on dimension count. The published evidence points to input resolution as the larger effect.</p>
<p>Jina CLIP v2 ablates image resolution separately, and the effect is much larger than anything in sections 3 or 4. On ViDoRe, moving from <strong>224 to 384 pixels lifts average NDCG@5 from 0.256 to 0.454</strong>. Raising resolution to 512 pixels improves it further; going from 512 to 768 costs 2.25x the image patches for a gain of 0.019, which is why the authors settle on 512.</p>
<p>Compare the two effects directly: cutting a well-trained MRL vector from 1024 to 512 dimensions costs a fraction of a point, while under-resolving the image costs a large share of retrieval quality. Preprocessing can destroy the information before vector width ever becomes the binding constraint.</p>
<p>This matters most for queries that hinge on a small local feature:</p>
<blockquote>
<p>"gas cooktop beneath a concealed rangehood" · "herringbone timber flooring" · "frameless shower with niche"</p>
</blockquote>
<p>Check what your API does here: Cohere documents that Embed v4 downsamples images above 2,458,624 pixels (1568×1568 square), while Gemini's embedding docs publish no equivalent threshold. One caveat: ViDoRe is document screenshots, not property photography, so do not transfer the specific 512px optimum. What transfers is the ordering of the two effects: fix resolution first, then choose a dimension.</p>
<p>One failure mode is unaffected by either setting: an embedding can correctly identify an image as a kitchen and still rank it top for <em>"white kitchen, black island"</em> when it is a black kitchen with a white island. Attribute binding is what COCO-style benchmarks underexpose.</p>
<h2>8. The corrected mental model</h2>
<p>Pulling the findings together:</p>
<ol>
<li><p><strong>Choose the model family first, then the dimension.</strong> Reversing the order risks picking a weaker model because it exposes more coordinates.</p>
</li>
<li><p><strong>Within an MRL model, 3-4x compression plausibly costs about a point.</strong> Nova and Jina agree on this from different labs and at different scales.</p>
</li>
<li><p><strong>The cliff is real, and it sits well below the headline width.</strong> Nova breaks at 256, Jina at 128.</p>
</li>
<li><p><strong>Fine-grained visual retrieval degrades faster than object-centric retrieval.</strong> ViDoRe falls fastest in Nova's table, and I would treat property photography as nearer that end.</p>
</li>
<li><p><strong>Resolution before dimension.</strong> It is the larger measured effect in the published evidence.</p>
</li>
<li><p><strong>Dimension and numeric precision are separate axes.</strong> Cohere exposes int8 and binary output; those compound with width rather than substituting for it, and conflating the two will confuse your results. It is a real axis, but it needs its own post.</p>
</li>
</ol>
<p>My current view on the operating regions — a hypothesis to test on your own corpus, not a set of guaranteed scores: 3072→1536 is rarely justifiable on quality alone; 1536/1024→768/512 is usually the attractive region; 384/256 is viable when memory-constrained but increasingly lossy on complex imagery; 128/64 is too aggressive without a reranker behind it.</p>
<p>That returns us to the 92 GB savings:</p>
<table>
<thead>
<tr>
<th>Dims</th>
<th>Bytes/vector</th>
<th>Raw vectors @ 10M</th>
<th>vs 3072</th>
</tr>
</thead>
<tbody><tr>
<td>3072</td>
<td>12 KiB</td>
<td>122.88 GB</td>
<td>100%</td>
</tr>
<tr>
<td>1536</td>
<td>6 KiB</td>
<td>61.44 GB</td>
<td>50%</td>
</tr>
<tr>
<td>1024</td>
<td>4 KiB</td>
<td>40.96 GB</td>
<td>33%</td>
</tr>
<tr>
<td>768</td>
<td>3 KiB</td>
<td>30.72 GB</td>
<td>25%</td>
</tr>
<tr>
<td>512</td>
<td>2 KiB</td>
<td>20.48 GB</td>
<td>17%</td>
</tr>
<tr>
<td>384</td>
<td>1.5 KiB</td>
<td>15.36 GB</td>
<td>12.5%</td>
</tr>
<tr>
<td>256</td>
<td>1 KiB</td>
<td>10.24 GB</td>
<td>8.3%</td>
</tr>
</tbody></table>
<p>Those figures are arithmetic on raw float32 payload only. A real ANN index also carries graph or product-quantisation structures, IDs and metadata, and latency has to be measured in your target vector database: dimension is a first-order effect on distance computation and memory traffic, but traversal, cache behaviour and I/O can dominate.</p>
<p>The conclusion the table supports: <strong>Gemini at 768 instead of 3072 dimensions saves roughly 92 GB of raw vectors on a 10M-photo catalogue, at one of Google's own recommended MRL sizes.</strong> That is why a sub-point quality loss deserves this much attention.</p>
<h2>9. Limits of this analysis</h2>
<p>Four limits I want on the record.</p>
<ul>
<li><p>Every number here comes from a <strong>general benchmark</strong>. None is a property-image benchmark.</p>
</li>
<li><p>Nova and Jina agreeing is suggestive, but it does not guarantee that <strong>Gemini's or Cohere's</strong> curves have the same shape. They may well not.</p>
</li>
<li><p>The 768 starting point for Gemini is <strong>extrapolated</strong> from Google's MRL training sizes plus Nova and Jina's behaviour. I could not find it demonstrated anywhere.</p>
</li>
<li><p>Cliff location is model-specific. You will have to find your own, and a few hundred well-judged domain queries will tell you more than another general benchmark will.</p>
</li>
</ul>
<p>That reverses the question we started with. The useful question was never how many dimensions a model has. It is how far down that model's own curve you can go before your queries degrade.</p>
<h2>Appendix A: Primary sources</h2>
<p>These are where the numbers in this post come from. I have noted which sections lean on each one so you can check the claims against the original.</p>
<ol>
<li><p><strong>Amazon Nova Multimodal Embeddings technical report.</strong> Amazon. <a href="https://cdn.amazon.science/ba/f2/d0af272848748a24ba6ba45af3a7/nova-mme-technical-report-14.pdf">cdn.amazon.science</a></p>
<p>Per-dimension text→image ablation at 3072 / 1024 / 384 / 256 (section 3), the prefix-truncation design (section 3), and the Cohere Embed 4 comparison run through Bedrock (section 6).</p>
</li>
<li><p><strong>jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images.</strong> Jina AI. <a href="https://arxiv.org/html/2412.08802v2">arXiv:2412.08802</a></p>
<p>Recall@5 across six MRL lengths (section 4) and the image-resolution ablation on ViDoRe (section 7).</p>
</li>
<li><p><strong>Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini.</strong> Google. <a href="https://arxiv.org/abs/2605.27295">arXiv:2605.27295</a></p>
<p>Native multimodal architecture, MRL training at 768- and 1536-d prefixes, and absolute text→image Recall@1 on DOCCI, TextCaps and MSCOCO (section 6).</p>
</li>
<li><p><strong>Embeddings, Gemini API documentation.</strong> Google AI for Developers. <a href="https://ai.google.dev/gemini-api/docs/embeddings">ai.google.dev</a></p>
<p>Supported range and recommended sizes (section 6), automatic re-normalisation at non-default sizes (section 5), and the Gemini Embedding 001 MTEB dimension table (section 6).</p>
</li>
<li><p><strong>Introduction to Embeddings at Cohere.</strong> Cohere documentation. <a href="https://docs.cohere.com/docs/embeddings">docs.cohere.com</a></p>
<p>Embed 4 dimension options (section 6), the image downsampling threshold (section 7), and int8 / binary output types (section 8).</p>
</li>
<li><p><strong>Embed: Secure AI Retrieval.</strong> Cohere product page. <a href="https://cohere.com/embed">cohere.com/embed</a></p>
<p>Embed 4's multimodal positioning and fused image-and-text vectors (section 6).</p>
</li>
</ol>
<h2>Appendix B: Background technique and benchmarks</h2>
<p>Useful if you want the original definition of a method or benchmark mentioned above. None of the post's figures come directly from these.</p>
<ul>
<li><p><strong>Matryoshka Representation Learning.</strong> Kusupati et al., 2022. <a href="https://arxiv.org/abs/2205.13147">arXiv:2205.13147</a>. The training method behind truncatable embeddings (section 5).</p>
</li>
<li><p><strong>ColPali: Efficient Document Retrieval with Vision Language Models.</strong> Faysse et al., 2024. <a href="https://arxiv.org/abs/2407.01449">arXiv:2407.01449</a>. Introduces the ViDoRe benchmark (sections 3 and 7).</p>
</li>
<li><p><strong>TextCaps: a Dataset for Image Captioning with Reading Comprehension.</strong> Sidorov et al., 2020. <a href="https://arxiv.org/abs/2003.12462">arXiv:2003.12462</a> (sections 3 and 6).</p>
</li>
<li><p><strong>Microsoft COCO: Common Objects in Context.</strong> Lin et al., 2014. <a href="https://arxiv.org/abs/1405.0312">arXiv:1405.0312</a> (sections 3 and 6).</p>
</li>
<li><p><strong>DOCCI: Descriptions of Connected and Contrasting Images.</strong> Onoe et al., 2024. <a href="https://arxiv.org/abs/2404.19753">arXiv:2404.19753</a> (section 6).</p>
</li>
<li><p><strong>Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset.</strong> Thapliyal et al., 2022. <a href="https://arxiv.org/abs/2205.12522">arXiv:2205.12522</a> (section 4).</p>
</li>
<li><p><strong>Towards Zero-shot Cross-lingual Image Retrieval.</strong> Aggarwal and Kale, 2020. <a href="https://arxiv.org/abs/2012.05107">arXiv:2012.05107</a>. Introduces XTD10 (section 4).</p>
</li>
<li><p><strong>MTEB: Massive Text Embedding Benchmark.</strong> Muennighoff et al., 2022. <a href="https://arxiv.org/abs/2210.07316">arXiv:2210.07316</a> (section 6).</p>
</li>
</ul>
]]></content:encoded></item></channel></rss>