verifier.org

BGE-M3

mit, three retrieval modes in one model, and quietly ageing

$0.010 / 1m tokens on DeepInfra#hybrid-retrieval-where-y#sparse-vectors-from-a-si

the most versatile open model here — dense, sparse and multi-vector output under mit — on a card that has barely moved since 2024.

BGE-M3 does something none of the others do: one model produces dense embeddings, sparse lexical weights and ColBERT-style multi-vector output. hybrid search normally means running a dense model alongside bm25 and reconciling two systems; this collapses that into one call, and for retrieval quality on mixed keyword-and-semantic queries it is a real architectural simplification.

the licence is mit, which is as permissive as this category gets, and it claims more than a hundred working languages across an 8,192-token context. it remains one of the most widely deployed open embedding models in production, and there is a matching bge-reranker-v2 family under the same licence.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
$0.010 / 1m tokens on DeepInfra
our verdict

the most versatile open model here — dense, sparse and multi-vector output under mit — on a card that has barely moved since 2024.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more embedding models

text-embedding-3-small

openai's cheaper embedding model, shortenable to fewer dimensions

#api#matryoshka

multilingual-e5

microsoft's open multilingual embedding family, widely used as a baseline

#open-weights#multilingual

ColBERT

late-interaction retrieval that scores token by token rather than one vector

#late-interaction#research

Qwen3-Embedding-8B

we checked this$0.010 / 1m tokens on DeepInfra

apache-2.0, best open multilingual quality, and a tenth of a cent hosted

#multilingual-retrieval-w#the-bill-has-to-be-small