verifier.org

gemini-embedding-2

one vector space for text, images, audio and video

$0.20 / 1m text tokens#retrieval-across-mixed-m

the only model here that puts text, images, audio and video in one shared space — at the highest text price in the ranking, and it just invalidated its own predecessor's vectors.

gemini-embedding-2 is Google's first multimodal embedding model, mapping text, images, video, audio and documents into a single space at over a hundred languages. that is a genuinely different capability from everything else here except Cohere, which does text and images but not audio or video. if your corpus is a mix of media and you want one index over all of it, this is the shortlist.

dimensions run from 128 to 3,072 with matryoshka support and recommended stops at 768, 1,536 and 3,072, so you can trade index size against quality deliberately. context is 8,192 tokens, four times what the older gemini-embedding-001 accepted.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
$0.20 / 1m text tokens
our verdict

the only model here that puts text, images, audio and video in one shared space — at the highest text price in the ranking, and it just invalidated its own predecessor's vectors.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more embedding models

text-embedding-3-small

openai's cheaper embedding model, shortenable to fewer dimensions

#api#matryoshka

multilingual-e5

microsoft's open multilingual embedding family, widely used as a baseline

#open-weights#multilingual

ColBERT

late-interaction retrieval that scores token by token rather than one vector

#late-interaction#research

Qwen3-Embedding-8B

we checked this$0.010 / 1m tokens on DeepInfra

apache-2.0, best open multilingual quality, and a tenth of a cent hosted

#multilingual-retrieval-w#the-bill-has-to-be-small