Static embeddings: a practical guide to model2vec
Sentence-transformer embeddings work well, but they are slow on a CPU. For a few thousand texts, the speed doesn't matter. When you embed millions of texts, or need embeddings within a tight latency budget, the compute cost and latency add up quickly.
model2vec reduces that cost. It distils a sentence-transformer into a static lookup table of token vectors. MinishLab, the team behind it, report that the distilled model is up to 500x faster and up to 50x smaller than the original. The distilled model is also less accurate, and the size of the drop depends on the task. This post explains how model2vec works, which model to pick, how it compares with transformers and older approaches, and which tasks I think it suits.
What model2vec is
model2vec is an open-source library from MinishLab, first released in 2024. Its only major runtime dependency is numpy.
Distillation has three steps:
Pass every token in the vocabulary through a teacher sentence-transformer. This gives one vector per token.
Reduce the vectors' dimensions with PCA (principal component analysis).
Down-weight common tokens using their estimated frequency (based on Zipf's law), so words like "the" contribute less to the final embedding.
To embed a text, the model tokenises it, looks up each token's vector, and averages the vectors. There is no attention layer, so it runs fast on a CPU and doesn't need a GPU.
Because each token's vector is fixed, a token gets the same vector in every context. The model also ignores word order.
from model2vec import StaticModel
model = StaticModel.from_pretrained("minishlab/potion-base-8M")
embeddings = model.encode(["The cat sat on the mat", "A kitten rested on the rug"])
You can also distil your own model from any sentence-transformer. MinishLab say this takes about 30 seconds on a CPU and needs no training data. You can also add your own domain vocabulary. An optional second step, Tokenlearn, trains the static vectors to match the teacher's sentence embeddings over a corpus. The potion models in the next section were trained with Tokenlearn.
Which model to pick
| Model | Params | Dims | Size | Use |
|---|---|---|---|---|
| potion-base-8M | 7.5M | 256 | 30 MB | General English, small & fast |
| potion-base-32M | 32.3M | 512 | 129 MB | Best general English quality |
| potion-retrieval-32M | 32.3M | 512 | 129 MB | English retrieval |
| potion-multilingual-128M | 128M | 256 | 512 MB | 101 languages |
| potion-code-16M-v2 | 16.2M | 256 | 32 MB | Code search |
The English models are distilled from bge-base-en-v1.5. The multilingual model is much larger because its vocabulary has about 500,000 tokens, and each token needs its own vector. The sizes are the model files on Hugging Face, stored as float32, except the code model, which is float16. MinishLab also publish smaller models: potion-base-2M (7.6 MB) and potion-base-4M (15 MB). The potion models replace the older M2V_* models.
I suggest starting with potion-base-8M, and moving to potion-base-32M if your evaluation shows a quality gap.
How it compares with sentence-transformers
MinishLab compare their models with all-MiniLM-L6-v2, which was released in 2021. Newer small models score much higher than MiniLM, so I've added two for comparison: bge-small-en-v1.5 and embeddinggemma-300m (released September 2025). The scores below are from the official MTEB(eng, v2) results (MTEB is the Massive Text Embedding Benchmark), averaged over task types.
| Task | potion-8M | potion-32M | MiniLM-L6 | bge-small | EmbeddingGemma |
|---|---|---|---|---|---|
| Params | 7.5M | 32M | 23M | 33M | 308M |
| Average | 51.08 | 52.13 | 55.93 | 59.93 | 65.11 |
| Classification | 70.34 | 71.70 | 69.25 | 76.56 | 87.55 |
| Clustering | 39.74 | 41.25 | 44.90 | 47.02 | 56.55 |
| STS | 72.91 | 73.93 | 78.95 | 81.28 | 83.61 |
| Retrieval | 31.11 | 32.67 | 42.92 | 53.86 | 55.69 |
The table shows four things:
On classification, potion matches or slightly beats MiniLM. bge-small scores about 5 points higher than potion-32M, and EmbeddingGemma about 16 points higher.
On clustering and STS (semantic textual similarity), potion-32M scores 4 to 17 points lower than the transformers.
On retrieval, potion-32M scores 10 to 23 points lower, which is the largest gap.
potion-base-32M keeps about 86% of its teacher's score (52.13 against 60.77 for
bge-base-en-v1.5).
model2vec scores lower than modern transformers on every task type. The reason to use it is speed and cost. MinishLab report up to 500x faster inference, and the smallest model is about 8 MB. The numpy-only runtime also gives faster start-up and smaller deployments.
Static embeddings can't represent meaning that depends on context. "Not good" ends up close to "good", word order has no effect, and a word like "bank" gets the same vector in every sentence.
model2vec and sentence-transformers are compatible. sentence-transformers can load model2vec models through its StaticEmbedding module. Tom Aarsen at Hugging Face has also used sentence-transformers to train related static models, such as static-retrieval-mrl-en-v1, with contrastive loss.
How it compares with other approaches
TF-IDF and BM25 are the fastest option and a strong baseline. They only match exact words, so they treat "car" and "vehicle" as unrelated.
GloVe and BPEmb are older static embeddings. Their MTEB averages are 45.82 and 41.74, against 51.08 for potion-8M. I think model2vec scores higher mainly because it uses subword tokens and learns from a modern teacher model, while GloVe learns from raw word co-occurrence counts.
WordLlama also builds static token vectors, but takes them from an LLM's token embeddings.
Small transformers such as MiniLM, bge-small and granite-embedding-small-english-r2 model context, but need far more compute per text. ONNX and int8 quantisation make them faster, but they are still much slower than a lookup table.
Small LLM-based embedders such as EmbeddingGemma-300m and Qwen3-Embedding-0.6B give the best quality at small sizes, with MTEB averages around 65. They have 10 to 25 times more parameters than MiniLM, so they're expensive to run on a CPU at high volume.
SetFit is a strong few-shot classifier, but it runs a transformer. In MinishLab's classification benchmark, a fine-tuned model2vec classifier was about 35x faster than SetFit on a CPU.
Embedding APIs give strong quality, but you pay per request, wait on network calls, and send your data to a third party.
| Approach | Speed on CPU | Quality | Uses context |
|---|---|---|---|
| TF-IDF / BM25 | Fastest | Low | No |
| GloVe / BPEmb | Very fast | Low | No |
| model2vec | Very fast | Medium | No |
| Small transformers | Slow | Good | Yes |
| LLM-based embedders | Very slow | Best (local) | Yes |
| Embedding APIs | Network-bound | Best | Yes |
Use cases
Text classification. Encode your texts and train a simple classifier on the embeddings:
from sklearn.linear_model import LogisticRegression
clf = LogisticRegression(max_iter=1000).fit(model.encode(train_texts), train_labels)
preds = clf.predict(model.encode(test_texts))
For better accuracy, use StaticModelForClassification (in the model2vec[train] extra), which fine-tunes the embeddings and the classifier together. Accuracy drops most on tasks where the label depends on negation or word order, such as some sentiment tasks.
Clustering. Encode the documents, then cluster the embeddings with KMeans or HDBSCAN. Embedding millions of documents is cheap enough to do on a laptop. The clusters will be less coherent than a transformer's: on MTEB clustering, the transformers above score 4 to 15 points higher than potion-32M.
from sklearn.cluster import KMeans
cluster_ids = KMeans(n_clusters=20).fit_predict(model.encode(docs))
Semantic similarity. The potion models return normalised vectors, so the dot product of two embeddings is their cosine similarity. This works well for near-duplicate detection and finding related texts. Because the model ignores word order, "dog bites man" and "man bites dog" get identical vectors.
a, b = model.encode(["The film was great", "I loved the movie"])
similarity = a @ b
Feature generation. Embeddings make good input features for gradient-boosted trees or other models. Encoding is cheap, so you can add a text column's embeddings to your existing tabular features without slowing the pipeline much:
import numpy as np
features = np.hstack([tabular_features, model.encode(df["description"].tolist())])
Retrieval. This post doesn't cover retrieval in depth. potion-retrieval-32M scores 35.06 on MTEB retrieval, against 42.92 for MiniLM and 53.86 for bge-small. I think it works as a fast first stage in front of a reranker, but it can't replace a strong retrieval model.
When not to use it
Don't use model2vec when your task depends on negation, word order or long-range context, or when retrieval quality is the main goal. For other tasks, I start with model2vec, evaluate it on my own data, and switch to a small transformer only where the quality gap matters.
Takeaways
Static embeddings trade some quality for a very large speed gain.
On classification, model2vec matches MiniLM, but modern small transformers score clearly higher.
If your workload is CPU-bound and high-volume, try model2vec first. It may be accurate enough for your task.
Sources
MinishLab/model2vec README and results
MTEB results, benchmark MTEB(eng, v2), averaged over task types
Train 400x faster static embedding models with Sentence Transformers (Tom Aarsen, Hugging Face)

