BGE-M3
mit, three retrieval modes in one model, and quietly ageing
the most versatile open model here — dense, sparse and multi-vector output under mit — on a card that has barely moved since 2024.
BGE-M3 does something none of the others do: one model produces dense embeddings, sparse lexical weights and ColBERT-style multi-vector output. hybrid search normally means running a dense model alongside bm25 and reconciling two systems; this collapses that into one call, and for retrieval quality on mixed keyword-and-semantic queries it is a real architectural simplification.
the licence is mit, which is as permissive as this category gets, and it claims more than a hundred working languages across an 8,192-token context. it remains one of the most widely deployed open embedding models in production, and there is a matching bge-reranker-v2 family under the same licence.
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- embedding models
- pricing
- $0.010 / 1m tokens on DeepInfra
- website
- huggingface.co
the most versatile open model here — dense, sparse and multi-vector output under mit — on a card that has barely moved since 2024.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.
more embedding models
text-embedding-3-small
openai's cheaper embedding model, shortenable to fewer dimensions
multilingual-e5
microsoft's open multilingual embedding family, widely used as a baseline