Nomic Embed Text v2
a mixture-of-experts embedder that only wakes up two thirds of itself
a clever, genuinely open, efficient multilingual model undone for most rag work by a 512-token context.
Nomic Embed Text v2 is a mixture-of-experts model — 475 million total parameters of which only 305 million activate on any given inference. that is an unusual design for an embedding model and it pays off exactly where you would hope: multilingual coverage of around a hundred languages at a compute cost closer to a small dense model than its total size suggests.
it is apache-2.0 on the weights, so there are no commercial conditions to work around, and it emits 768 dimensions truncatable to 256 for roughly a threefold storage reduction. Nomic has a good record of publishing training data and methodology alongside its models, which is rarer than it should be.
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- embedding models
- pricing
- free — self-host
- website
- huggingface.co
a clever, genuinely open, efficient multilingual model undone for most rag work by a 512-token context.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.
more embedding models
text-embedding-3-small
openai's cheaper embedding model, shortenable to fewer dimensions
multilingual-e5
microsoft's open multilingual embedding family, widely used as a baseline