verifier.org

Qwen3-Embedding-8B alternatives

11 tools we tested head to head against Qwen3-Embedding-8B, ranked — and what each one actually does differently.

last reviewed 21 aug 2026 · from our best 12 embedding models ·list curated by Onur Ozcanxin

first — what you'd be leaving

Qwen3-Embedding-8B ranks #1 of 12 in our embedding models testing. apache-2.0, best open multilingual quality, and a tenth of a cent hosted.

89/100

the best combination of open licence, multilingual quality and truncatable dimensions in the category — provided you can afford to serve an 8b model or are happy renting one.

why people look for an alternative
  • 8b parameters — the heaviest model here to self-host
  • 4,096 native dimensions is a large index if you do not truncate
  • text only, no multimodal path

stay with Qwen3-Embedding-8B if apache-2.0 on code and weights — commercially safe without conditions is the thing you care about most — nothing below beats it on that.

the short version
best alternativegemini-embedding-2retrieval across mixed media where text-only embeddings cannot answer the question86/100best free optionStellaenglish-only retrieval where you want mit weights and full control of index size68/100
advertisement
  1. 1

    gemini-embedding-2

    #2 in embedding models · one vector space for text, images, audio and video

    86/100

    verdictthe only model here that puts text, images, audio and video in one shared space — at the highest text price in the ranking, and it just invalidated its own predecessor's vectors.

    gemini-embedding-2 vs Qwen3-Embedding-8B
     Qwen3-Embedding-8Bgemini-embedding-2
    price$0.010 / 1m tokens on DeepInfra$0.20 / 1m text tokens
    free tieryesyes
    price / 1m$0.010 hosted, free self-hosted$0.20 text, $0.10 batch
    dimensions4,096, truncatable to 32128 to 3,072
    context32,768 tokens8,192 tokens
    licenceApache-2.0proprietary api
    multimodalno — text onlyyes — text, image, audio, video

    switch forretrieval across mixed media where text-only embeddings cannot answer the question

    pros
    • +text, images, audio, video and documents in one shared vector space
    • +100+ languages, with matryoshka dimensions from 128 to 3,072
    • +batch mode halves the text rate to $0.10 per million
    • +8,192-token context, four times its own predecessor
    cons
    • $0.20 per million is the highest text rate in the ranking
    • audio costs 32x and video 60x the text rate
    • vectors are incompatible with gemini-embedding-001 — upgrading means re-embedding everything
    • no MTEB score published for this version yet
  2. 2

    text-embedding-3-large

    #3 in embedding models · the default everyone reaches for, and it is showing its age

    83/100

    verdictstill a solid general-purpose embedding with the most flexible dimension control here — on an 8k context and a model that has not been updated since january 2024.

    text-embedding-3-large vs Qwen3-Embedding-8B
     Qwen3-Embedding-8Btext-embedding-3-large
    price$0.010 / 1m tokens on DeepInfra$0.13 / 1m tokens
    free tieryesno
    price / 1m$0.010 hosted, free self-hosted$0.13
    dimensions4,096, truncatable to 323,072, truncatable to any size
    context32,768 tokens8,192 tokens
    licenceApache-2.0proprietary api
    multimodalno — text onlyno — text only

    switch forteams already on the OpenAI api who want a known quantity and no new vendor

    pros
    • +arbitrary dimension truncation via the dimensions parameter, not a fixed list
    • +no new vendor if you are already on the OpenAI api
    • +extremely well documented and widely supported by every framework
    • +the cheap sibling, text-embedding-3-small, is $0.02 per million
    cons
    • 8,192-token context is joint smallest in the ranking
    • no replacement shipped since january 2024
    • no multilingual claim, no multimodal support, no reranker sibling
    • quotes an MTEB score without naming the board version
  3. 3

    voyage-4-large

    #4 in embedding models · cheaper than OpenAI, four times the context, and a reranker to match

    82/100

    verdictbeats the default on price, context and reranking, from a smaller vendor that publishes less about how it performs.

    voyage-4-large vs Qwen3-Embedding-8B
     Qwen3-Embedding-8Bvoyage-4-large
    price$0.010 / 1m tokens on DeepInfra$0.12 / 1m tokens
    free tieryesyes
    price / 1m$0.010 hosted, free self-hosted$0.12
    dimensions4,096, truncatable to 321,024 default; 256 to 2,048
    context32,768 tokens32,000 tokens
    licenceApache-2.0proprietary api
    multimodalno — text onlyunconfirmed for this model

    switch forretrieval-focused teams who want a matched embedding and reranking pair from one vendor

    pros
    • +$0.12 per million — cheaper than OpenAI and it cut price on the version bump
    • +32,000-token context, four times OpenAI's
    • +matched rerankers at $0.05 and $0.02 per million
    • +dimension options are cross-compatible within the 4 series
    cons
    • no MTEB score published on its own pages
    • no language count or parameter count disclosed
    • multimodal support for this specific model is unconfirmed
    • smaller vendor than the three above it
    advertisement
  4. 4

    Cohere Embed v4

    #5 in embedding models · 128k of context, text and pdfs in one space, and no published price

    78/100

    verdictby far the longest context here and genuine text-plus-pdf embedding — from the only vendor in this ranking that will not tell you what an api call costs.

    Cohere Embed v4 vs Qwen3-Embedding-8B
     Qwen3-Embedding-8BCohere Embed v4
    price$0.010 / 1m tokens on DeepInfranot published per token
    free tieryesyes
    price / 1m$0.010 hosted, free self-hostednot published
    dimensions4,096, truncatable to 32256 / 512 / 1,024 / 1,536
    context32,768 tokens128,000 tokens
    licenceApache-2.0proprietary api
    multimodalno — text onlyyes — text, images, pdfs

    switch forlong-document and pdf retrieval where a 128k window removes the chunking problem

    pros
    • +128,000-token context — fifteen times OpenAI's, the longest here by far
    • +text, images and pdfs embedded into one space
    • +matching Rerank 3.5 and Rerank 4 models
    • +four selectable output dimensions from 256 to 1,536
    cons
    • no per-token api price published anywhere on cohere.com
    • only dedicated-instance pricing is public, from $2,500 a month
    • no MTEB score on its own pages, and third-party figures disagree
    • no language count stated for v4.0 specifically
  5. 5

    EmbeddingGemma

    #6 in embedding models · the cheapest way to embed anything, if the gemma terms suit you

    75/100

    verdict300m parameters that run on a laptop at a fiftieth of Gemini's price — held back by a licence that is commercially usable but not open source.

    EmbeddingGemma vs Qwen3-Embedding-8B
     Qwen3-Embedding-8BEmbeddingGemma
    price$0.010 / 1m tokens on DeepInfra$0.002 / 1m tokens on DeepInfra
    free tieryesyes
    price / 1m$0.010 hosted, free self-hosted$0.002 hosted, free self-hosted
    dimensions4,096, truncatable to 32768, truncatable to 128
    context32,768 tokens2,048 tokens
    licenceApache-2.0Gemma Terms of Use
    multimodalno — text onlyno — text only

    switch foron-device and high-volume embedding where size and cost dominate everything else

    pros
    • +$0.002 per million hosted — the cheapest rate in the ranking
    • +300m parameters, genuinely runs on a laptop or phone
    • +matryoshka truncation to 512, 256 or 128 dimensions
    • +publishes MTEB scores with the board version named
    cons
    • Gemma Terms of Use, not an osi licence — prohibited-use policy attached
    • redistribution obliges you to pass the same restrictions downstream
    • 2,048-token context is among the shortest here
    • quality sits below the larger open models, as expected at 300m
  6. 6

    BGE-M3

    #7 in embedding models · mit, three retrieval modes in one model, and quietly ageing

    73/100

    verdictthe most versatile open model here — dense, sparse and multi-vector output under mit — on a card that has barely moved since 2024.

    BGE-M3 vs Qwen3-Embedding-8B
     Qwen3-Embedding-8BBGE-M3
    price$0.010 / 1m tokens on DeepInfra$0.010 / 1m tokens on DeepInfra
    free tieryesyes
    price / 1m$0.010 hosted, free self-hosted$0.010 hosted, free self-hosted
    dimensions4,096, truncatable to 321,024, fixed
    context32,768 tokens8,192 tokens
    licenceApache-2.0MIT
    multimodalno — text onlyno — text only

    switch forhybrid retrieval where you want dense and sparse vectors from a single model

    pros
    • +mit licensed — the most permissive terms in the ranking
    • +dense, sparse and multi-vector retrieval from one model
    • +100+ languages at 8,192 tokens of context
    • +matching bge-reranker-v2 family under the same licence
    cons
    • fixed 1,024 dimensions with no matryoshka truncation
    • card essentially unchanged since july 2024
    • NVIDIA NIM is deprecating its hosted endpoint in august 2026
    • no MTEB score with an identifiable board version
  7. 7

    Stella

    #8 in embedding models · mit weights with dimensions from 256 to 8,192, english only

    68/100

    verdictthe widest dimension range in the category under an unrestricted mit licence — english-only, short-context, and with nobody hosting it for you.

    Stella vs Qwen3-Embedding-8B
     Qwen3-Embedding-8BStella
    price$0.010 / 1m tokens on DeepInfrafree — self-host only
    free tieryesyes
    price / 1m$0.010 hosted, free self-hostedfree — self-host only
    dimensions4,096, truncatable to 32256 to 8,192
    context32,768 tokens512 tokens recommended
    licenceApache-2.0MIT
    multimodalno — text onlyno — text only

    switch forenglish-only retrieval where you want mit weights and full control of index size

    pros
    • +mit licensed with no usage conditions whatsoever
    • +eight matryoshka dimension options from 256 to 8,192
    • +the authors state where quality plateaus, rather than leaving you to test
    • +1.5b parameters — far lighter to serve than Qwen3-8B
    cons
    • english only
    • 512-token recommended input, the shortest here
    • no hosted option we could price — self-host or nothing
    • licence comes from card metadata; the repo licence file 404s
  8. 8

    Nomic Embed Text v2

    #9 in embedding models · a mixture-of-experts embedder that only wakes up two thirds of itself

    66/100

    verdicta clever, genuinely open, efficient multilingual model undone for most rag work by a 512-token context.

    Nomic Embed Text v2 vs Qwen3-Embedding-8B
     Qwen3-Embedding-8BNomic Embed Text v2
    price$0.010 / 1m tokens on DeepInfrafree — self-host
    free tieryesyes
    price / 1m$0.010 hosted, free self-hostedfree — self-host
    dimensions4,096, truncatable to 32768, truncatable to 256
    context32,768 tokens512 tokens
    licenceApache-2.0Apache-2.0
    multimodalno — text onlyno — text only

    switch forself-hosted multilingual embedding on constrained compute

    pros
    • +apache-2.0 weights with no commercial conditions
    • +mixture-of-experts design activates only 305m of 475m parameters
    • +around 100 languages from a very small model
    • +matryoshka truncation from 768 to 256 dimensions
    cons
    • 512-token context makes it unsuitable for document retrieval
    • no hosted list price confirmable from a primary source
    • no MTEB score published on the model card
    • no reranker sibling
  9. 9

    Mistral Embed

    #10 in embedding models · a flat price, a tidy api, and almost no published specification

    62/100

    verdictcheaper than OpenAI and easy to adopt if you are already a Mistral customer — from a model that publishes neither its dimensions nor its benchmark scores.

    Mistral Embed vs Qwen3-Embedding-8B
     Qwen3-Embedding-8BMistral Embed
    price$0.010 / 1m tokens on DeepInfra$0.10 / 1m tokens
    free tieryesyes
    price / 1m$0.010 hosted, free self-hosted$0.10
    dimensions4,096, truncatable to 32not published by vendor
    context32,768 tokens8,000 tokens
    licenceApache-2.0proprietary api
    multimodalno — text onlyno — text and code

    switch forteams already building on Mistral who want embeddings on the same bill

    pros
    • +$0.10 per million, cheaper than OpenAI's flagship
    • +no new vendor if you are already using the Mistral api
    • +flat pricing with no tiers or dimension-based surcharges
    • +a separate Codestral Embed exists for code retrieval
    cons
    • output dimensions are not published by the vendor
    • no parameter count, language count or MTEB score published
    • dates to december 2023 — the oldest model in this ranking
    • no weights, no self-host path, no reranker sibling
  10. 10

    Jina Embeddings v5

    #11 in embedding models · excellent quality per parameter, on weights you may not use commercially

    58/100

    verdicta genuinely strong sub-1b multilingual model that most readers of this page cannot legally deploy — and the licence is not what its reputation suggests.

    Jina Embeddings v5 vs Qwen3-Embedding-8B
     Qwen3-Embedding-8BJina Embeddings v5
    price$0.010 / 1m tokens on DeepInfranot published per token
    free tieryesyes
    price / 1m$0.010 hosted, free self-hostednot published
    dimensions4,096, truncatable to 321,024, truncatable to 32
    context32,768 tokens32,768 tokens
    licenceApache-2.0CC-BY-NC-4.0 — non-commercial
    multimodalno — text onlyno — v5-omni is a separate model

    switch forresearch and evaluation work where the non-commercial terms are not a problem

    pros
    • +677m parameters with 32,768 tokens of context
    • +matryoshka truncation from 1,024 down to 32 dimensions
    • +93 languages claimed from a sub-billion-parameter model
    • +a separate v5-omni line covers images, audio, video and pdfs
    cons
    • weights are CC-BY-NC-4.0 — commercial use prohibited without a separate licence
    • v4 was non-commercial too, under a qwen research licence
    • no per-token price published for the hosted api
    • the v4-to-v5 split moved multimodality to a different model line
+ 1 more tested, not detailed here
we ranked 12 embedding models in total. the 1 that didn't make this page are written up in the full ranking →

how these were compared

every tool on this page went through the same test as Qwen3-Embedding-8B — same tasks, same order, scored the same way. the comparison tables are the figures from that testing, not vendor spec sheets.

the embedding models test in full →
was this useful?