Introduction
E5 (EmbEddings from bidirEctional Encoder rEpresentations) is a family of open-source text embedding models developed by Microsoft, distinguished by a training approach built around massive-scale weakly-supervised contrastive pretraining, followed by fine-tuning on labeled data. Like BGE (covered in the previous topic), E5 models are open-weight, self-hostable, and consistently rank among the top performers on standard embedding benchmarks — and they share a similar, distinctive convention: using explicit prefixes to distinguish between queries and documents/passages.
E5 models round out this section's tour of leading embedding options, alongside the API-based convenience of OpenAI Embeddings and the open-source strength of Sentence Transformers and BGE — giving a well-rounded picture of the choices available when selecting an embedding model for a real application.
Why Do E5 Models Matter?
E5 models help to:
- Provide open-source, self-hostable embeddings with strong benchmark performance
- Offer multiple model sizes, balancing embedding quality against speed and resource requirements
- Support strong multilingual capabilities through the dedicated multilingual-E5 variants
- Use explicit query/passage prefixes to help the model distinguish between retrieval roles
- Integrate directly with the
sentence-transformerslibrary, like the other open-source models in this section - Represent Microsoft's contribution to the actively evolving open embedding model landscape
Using E5 Models via Sentence Transformers
Unlike BGE's instruction prefix (which is recommended
specifically for queries only), E5 models require explicit
prefixes on BOTH the query ("query: ") and the passage/document
("passage: ") side — omitting these prefixes, or using them
inconsistently, can noticeably hurt E5's retrieval performance,
since the model was specifically trained expecting this format.Why Prefixes Matter for E5 Specifically
E5 was trained using massive-scale contrastive learning, where
the model repeatedly learns to distinguish matching query-passage
pairs from non-matching ones. The "query: " and "passage: "
prefixes were part of this training process from the start —
they aren't just a suggested convention, but an integral part
of how the model learned to represent these two distinct roles
differently, even when the underlying text might otherwise look
similar.
This is a stricter, more required convention than BGE's
recommended-but-more-flexible query-only instruction prefix,
making it especially important to follow E5's documentation
exactly for expected performance.E5 Model Size Variants
| Variant | General Characteristics |
|---|---|
| e5-small | Smallest, fastest, lowest resource requirements |
| e5-base | Balanced tradeoff between speed and embedding quality |
| e5-large | Highest quality among standard sizes, higher resource cost |
| multilingual-e5 (small/base/large) | Same size tiers, trained specifically for strong multilingual performance |
A Practical Retrieval Example
This directly builds on the RAG-style pattern covered in the OpenAI Embeddings topic, just using E5's specific self-hosted model and required prefix convention instead.
E5 Models vs BGE Models
| Aspect | E5 Models | BGE Models |
|---|---|---|
| Developer | Microsoft | BAAI |
| Prefix Convention | Required on BOTH query and passage sides | Recommended, typically query-side only |
| Training Approach | Massive-scale contrastive pretraining + fine-tuning | Large-scale pretraining with instruction-aware fine-tuning |
| Multilingual Variant | multilingual-e5 | bge-m3 (also multi-granularity, multi-functionality) |
| Benchmark Performance | Frequently near the top of leaderboards like MTEB | Frequently near the top of leaderboards like MTEB |
In practice, BGE and E5 are close competitors, often trading
places at the top of embedding benchmark leaderboards as new
versions of each are released — the "best" choice between them
frequently depends on your specific domain, language needs, and
direct benchmarking against your own actual data, rather than
a fixed, universal answer.E5 Models vs OpenAI Embeddings
| Aspect | E5 Models | OpenAI Embeddings |
|---|---|---|
| Access | Self-hosted, open-weight | API-based, closed-weight |
| Cost Structure | Infrastructure/compute cost | Per-request/per-token pricing |
| Data Privacy | Full control, data never leaves your infrastructure | Data sent to OpenAI's servers |
| Setup Effort | Requires managing model hosting | Minimal — just an API call |
| Customization | Can be fine-tuned further if needed | Limited to what OpenAI provides |
Key Properties of E5 Models
- E5 models are open-source embedding models developed by Microsoft, trained via large-scale contrastive learning.
- They require explicit "query: " and "passage: " prefixes on inputs, unlike BGE's more flexible convention.
- Multiple size variants (small, base, large) and dedicated multilingual versions are available.
- E5 models are accessed through the same
sentence-transformerslibrary used throughout this section. - E5 and BGE frequently compete closely at the top of standard embedding benchmarks like MTEB.
Where Are E5 Models Used?
| Field | Application |
|---|---|
| RAG-Based Applications | Self-hosted document retrieval for LLM context |
| Multilingual Search Systems | Leveraging multilingual-E5 for cross-language retrieval |
| Enterprise Document Search | Self-hosted, privacy-preserving semantic search |
| Academic and Research Benchmarking | A common strong baseline alongside BGE in embedding research |
| Cost-Sensitive High-Volume Applications | Avoiding per-request API costs at scale via self-hosting |
Advantages
- Open-source and self-hostable, avoiding per-request API costs and data privacy concerns
- Strong, benchmark-competitive performance, frequently rivaling or exceeding BGE and other alternatives
- Multiple size variants allow tuning the quality-vs-speed tradeoff for a given application
- Dedicated multilingual variants support strong cross-language retrieval performance
- Compatible with the widely-used and well-documented
sentence-transformerslibrary
Limitations
- Requires self-hosting infrastructure, unlike a fully managed API-based option
- The required query/passage prefix convention is easy to forget or apply inconsistently, hurting performance
- Larger E5 variants require meaningful compute resources for good performance
- Benchmark leaderboard rankings shift over time as new models (including newer E5 versions) are released
- Requires validating performance on your specific domain, not just relying on general benchmark rankings
Real-World Examples
| Application | E5 Model Use |
|---|---|
| Self-Hosted RAG Systems | Using e5-base or e5-large for high-quality document retrieval |
| Multilingual Search Platforms | Using multilingual-e5 for cross-language semantic search |
| Enterprise Knowledge Bases | Self-hosted E5 embeddings for internal, privacy-sensitive document search |
| Academic Research Benchmarking | Comparing new embedding techniques against E5 as a strong baseline |
| Cost-Optimized Production Search | High-volume applications avoiding per-request API costs via self-hosting |
Best Practices
- Always apply the "query: " and "passage: " prefixes exactly as required — this isn't optional for E5 models.
- Use
normalize_embeddings=Truewhen encoding, consistent with E5's designed convention. - Choose an E5 size variant based on your specific latency, cost, and quality requirements.
- Use a multilingual-E5 variant specifically when cross-language retrieval is needed.
- Benchmark E5 against BGE and other alternatives on your own domain-specific data before committing to one.
Interview Tip
A common interview question is:
"Why is it important to use the correct query/passage prefixes when working with E5 models, and what happens if you don't?"
A strong answer is:
E5 models were trained from the start using explicit "query: " and "passage: " prefixes as part of their massive-scale contrastive learning process, so the model learned to represent these two roles differently based partly on that prefix signal. If you omit the prefixes or apply them inconsistently — for example, forgetting the passage prefix while still using the query prefix — the resulting embeddings won't match the format the model was actually trained and evaluated on, which can noticeably degrade retrieval accuracy, since the model may no longer clearly distinguish which piece of text is the query and which is the document being searched.
Explaining that the prefixes are trained-in, not just cosmetic, makes your answer stronger.
Conclusion
E5 models provide another strong, open-source, self-hostable embedding option, distinguished by their explicit and required query/passage prefix convention rooted directly in how the models were trained. With OpenAI Embeddings, Sentence Transformers, BGE Models, and now E5 Models all covered, this completes the full Embedding Models section, giving a well-rounded view of both API-based and open-source options for converting text into the vector representations that power semantic search, RAG, and similarity comparison throughout modern AI applications.