gemini-embedding-2
one vector space for text, images, audio and video
the only model here that puts text, images, audio and video in one shared space — at the highest text price in the ranking, and it just invalidated its own predecessor's vectors.
gemini-embedding-2 is Google's first multimodal embedding model, mapping text, images, video, audio and documents into a single space at over a hundred languages. that is a genuinely different capability from everything else here except Cohere, which does text and images but not audio or video. if your corpus is a mix of media and you want one index over all of it, this is the shortlist.
dimensions run from 128 to 3,072 with matryoshka support and recommended stops at 768, 1,536 and 3,072, so you can trade index size against quality deliberately. context is 8,192 tokens, four times what the older gemini-embedding-001 accepted.
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- embedding models
- pricing
- $0.20 / 1m text tokens
- website
- ai.google.dev
the only model here that puts text, images, audio and video in one shared space — at the highest text price in the ranking, and it just invalidated its own predecessor's vectors.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.
more embedding models
text-embedding-3-small
openai's cheaper embedding model, shortenable to fewer dimensions
multilingual-e5
microsoft's open multilingual embedding family, widely used as a baseline