Cohere has released Embed 5, a new embedding model family. It targets enterprise search, RAG, and agentic retrieval. The model family ships in 2 tiers. Embed 5 Pro targets maximum retrieval quality. Embed 5 Fast targets latency and cost on the live query path. Both accept text, images, and fused text plus image inputs. Both cover 100+ languages and read up to 128K tokens. The key design choice: Pro and Fast share 1 embedding space. You can index with one and query with the other.
Is it deployable today? Yes, both tiers are generally available on the Cohere API and Model Vault, Microsoft Foundry, and Amazon SageMaker. Private VPC or on-prem serving runs through vLLM.
What Cohere Shipped
The API model IDs are embed-v5.0-pro and embed-v5.0-fast, per Cohere’s model docs. Both output 2048, 1536, 1024, 768, 512, or 256 dimensions, with 2048 as default. Embeddings come back as float, int8, or binary. Pro costs $0.12 per 1M text tokens. Fast costs $0.08. Image inputs cost $0.40 per 1M tokens on both.
Embed 5 can embed a page image directly. It can also fuse an image with its metadata into a single vector. That is important for scanned pages, slide decks, schematics, and charts, where text extraction drops information.
Pro and Fast: One Index, Two Query Paths
Cohere tested every corpus and query pairing across 40 development datasets. Normalized to Pro plus Pro at 100, a Pro index queried with Fast scored 98.4. An all-Fast setup scored 96.6. Cohere’s recommended pattern is to index with Pro and query with Fast. One constraint: both sides must use the same output dimension.
The split targets agentic workloads. An agent may issue dozens of searches per task, and query latency compounds. Cohere team reports Fast processed 377.3 documents per second versus 159.7 for Pro.
Benchmarks
On ViDoRe V3, Embed 5 Pro averages 85.8, an 8.8-point gain over Embed 4. Fast averages 84.5. Voyage 4 Large scores 83.7, Gemini Embedding 2 scores 83.2, and OpenAI text-embedding-3-large scores 75.5. On Cohere’s parsed-PDF suite, Pro leads at 84.8 against Voyage 4 Large at 83.6.
Finance is the strongest showing. Pro ranks first on FinanceBench (80.1), FinQA (90.0), and ViDoRe V3 Finance (85.0). Fast ranks second on all 3.
Multilingual results are mixed. Pro leads the European-language average at 77. However, Gemini Embedding 2 beats Pro on 9 of 10 further languages in Cohere’s own results table. Those include Japanese, Arabic, Hindi, and Telugu.
One important thing to note. Most numbers use RCP-nDCG@10, a new Cohere metric. It reorders a fixed candidate set, so it measures reranking quality more than first-stage retrieval. Cohere published the evaluation code, but independent replication is still pending.
Storage Costs at Scale
Embed 5 uses Matryoshka representation learning plus lower-precision outputs. A 2048-dim float32 vector needs 8 KB. A 1024-dim int8 vector needs 1 KB. A 256-dim binary vector needs 32 bytes. Across 100M chunks, raw storage drops from about 819 GB to 3.2 GB. Cohere recommends 1024-dim int8 as the default, citing near-full-precision quality.
Interactive Explainer: How Embed 5 Works
Cohere Embed 5 · Interactive explainer1 / 6
Embed 5 vs Closest Competitors
| Feature | Cohere Embed 5 Pro | Cohere Embed 5 Fast | Voyage 4 Large | Gemini Embedding 2 | OpenAI text-embedding-3-large |
|---|---|---|---|---|---|
| Max input | 128K tokens | 128K tokens | 32,000 tokens | 8,192 tokens | 8,191 tokens |
| Input types | Text, image, fused text + image | Text, image, fused text + image | Text | Text, image, video, audio, PDF | Text |
| Output dimensions | 256 to 2048 (6 sizes) | 256 to 2048 (6 sizes) | 256, 512, 1024, 2048 | 128 to 3072 | Up to 3072 (shortenable) |
| Output formats | float, int8, binary | float, int8, binary | float, int8, uint8, binary, ubinary | float | float |
| Languages | 100+ | 100+ | Multilingual (count not published) | 100+ | Multilingual (count not published) |
| Shared space across tiers | Yes (with Fast) | Yes (with Pro) | Yes (Voyage 4 series) | No sibling tier | No sibling tier |
| Price per 1M text tokens | $0.12 | $0.08 | $0.12 | $0.20 | $0.13 |
| Private / self-hosted | Yes (VPC or on-prem via vLLM) | Yes (VPC or on-prem via vLLM) | Via AWS Marketplace model package | No (Gemini API, Vertex AI) | No (API only) |
| ViDoRe V3 avg (Cohere-reported) | 85.8 | 84.5 | 83.7 | 83.2 | 75.5 |
Sources: Cohere blog, Cohere docs, Voyage docs, Google Gemini docs, OpenAI docs. Verified October 1, 2026. Benchmark scores come from Cohere and use its RCP-nDCG@10 metric.
Key Takeaways
- Embed 5 Pro scores 85.8 on ViDoRe V3, ahead of Voyage 4 Large and Gemini Embedding 2.
- Embed 5 Fast costs $0.08 per 1M text tokens and runs 2.4x Pro’s document throughput.
- Pro and Fast share 1 vector space: index with Pro, query with Fast, no re-indexing.
- 128K-token context, 100+ languages, and multimodal inputs on both tiers.
- 256-dim binary vectors take 32 bytes, a 256x cut versus 2048-dim float32.
FAQ
- What is Cohere Embed 5? A multimodal, multilingual embedding model family from Cohere, released September 30, 2026, in Pro and Fast tiers.
- How much does Embed 5 cost? Pro costs $0.12 and Fast costs $0.08 per 1M text tokens. Images cost $0.40 per 1M tokens on both.
- Can I mix Pro and Fast embeddings? Yes. They share 1 embedding space, provided both use the same output dimension.
Check out the technical details, product page, and announcement on X. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.
Credit: Source link

























