verifier.org

Stella

mit weights with dimensions from 256 to 8,192, english only

free — self-host only#english-only-retrieval-w#full-control-of-index-si

the widest dimension range in the category under an unrestricted mit licence — english-only, short-context, and with nobody hosting it for you.

Stella is 1.5 billion parameters on a Qwen2 backbone, mit licensed, with matryoshka dimensions at 256, 512, 768, 1,024, 2,048, 4,096, 6,144 and 8,192. that is the widest range on offer here, and the authors note that 1,024 retains nearly all the quality of 8,192 — which is a more useful statement than most vendors make, because it tells you where the knee is instead of leaving you to find it.

mit means no conditions at all: no prohibited-use policy like EmbeddingGemma, no non-commercial clause like Jina or NV-Embed, no copyleft. for a team that needs to embed a model inside a shipped product, that matters more than a point of benchmark score.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
free — self-host only
our verdict

the widest dimension range in the category under an unrestricted mit licence — english-only, short-context, and with nobody hosting it for you.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more embedding models

text-embedding-3-small

openai's cheaper embedding model, shortenable to fewer dimensions

#api#matryoshka

multilingual-e5

microsoft's open multilingual embedding family, widely used as a baseline

#open-weights#multilingual

ColBERT

late-interaction retrieval that scores token by token rather than one vector

#late-interaction#research

Qwen3-Embedding-8B

we checked this$0.010 / 1m tokens on DeepInfra

apache-2.0, best open multilingual quality, and a tenth of a cent hosted

#multilingual-retrieval-w#the-bill-has-to-be-small