verifier.org

Nomic Embed Text v2

a mixture-of-experts embedder that only wakes up two thirds of itself

free — self-host#self-hosted-multilingual

a clever, genuinely open, efficient multilingual model undone for most rag work by a 512-token context.

Nomic Embed Text v2 is a mixture-of-experts model — 475 million total parameters of which only 305 million activate on any given inference. that is an unusual design for an embedding model and it pays off exactly where you would hope: multilingual coverage of around a hundred languages at a compute cost closer to a small dense model than its total size suggests.

it is apache-2.0 on the weights, so there are no commercial conditions to work around, and it emits 768 dimensions truncatable to 256 for roughly a threefold storage reduction. Nomic has a good record of publishing training data and methodology alongside its models, which is rarer than it should be.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
free — self-host
our verdict

a clever, genuinely open, efficient multilingual model undone for most rag work by a 512-token context.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more embedding models

text-embedding-3-small

openai's cheaper embedding model, shortenable to fewer dimensions

#api#matryoshka

multilingual-e5

microsoft's open multilingual embedding family, widely used as a baseline

#open-weights#multilingual

ColBERT

late-interaction retrieval that scores token by token rather than one vector

#late-interaction#research

Qwen3-Embedding-8B

we checked this$0.010 / 1m tokens on DeepInfra

apache-2.0, best open multilingual quality, and a tenth of a cent hosted

#multilingual-retrieval-w#the-bill-has-to-be-small