Qwen3-Embedding-8B
apache-2.0, best open multilingual quality, and a tenth of a cent hosted
the best combination of open licence, multilingual quality and truncatable dimensions in the category — provided you can afford to serve an 8b model or are happy renting one.
Qwen3-Embedding-8B is apache-2.0 on both code and weights, which puts it in the small group here you can ship commercially without reading anything twice. it claims over a hundred languages, takes 32,768 tokens of context — four times what OpenAI or Gemini accept — and emits 4,096 dimensions that matryoshka truncation can take all the way down to 32, so the storage cost is yours to choose rather than the model's to impose.
on quality it publishes the strongest open numbers here, and unusually it labels them: 75.22 on MTEB english v2 and 70.58 on the multilingual board, per its own card. we still would not buy on that alone, but a vendor that names the board version is doing better than most of this field.
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- embedding models
- pricing
- $0.010 / 1m tokens on DeepInfra
- website
- huggingface.co
the best combination of open licence, multilingual quality and truncatable dimensions in the category — provided you can afford to serve an 8b model or are happy renting one.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.
more embedding models
text-embedding-3-small
openai's cheaper embedding model, shortenable to fewer dimensions
multilingual-e5
microsoft's open multilingual embedding family, widely used as a baseline