Introduction

BGE (BAAI General Embedding) is a family of open-source text embedding models developed by the Beijing Academy of Artificial Intelligence (BAAI), known for achieving consistently strong results on embedding benchmarks while remaining fully open-weight and self-hostable. Alongside Sentence Transformers (covered in the previous topic), BGE models represent another leading option in the open-source embedding space, frequently appearing at or near the top of leaderboards like MTEB (the Massive Text Embedding Benchmark) used to compare embedding model quality across a wide range of tasks.

BGE models are typically used through the same sentence-transformers library covered in the previous topic, since they follow compatible conventions — meaning switching from a standard Sentence Transformers model to a BGE model often requires little more than changing the model name in your code.

Why Do BGE Models Matter?

BGE models help to:

  • Provide open-source, self-hostable embeddings with consistently strong benchmark performance
  • Offer multiple model sizes, balancing embedding quality against speed and resource requirements
  • Support strong multilingual capabilities in several BGE model variants
  • Include specialized instruction-aware variants for improved retrieval performance
  • Integrate directly with the widely-used sentence-transformers library and ecosystem
  • Serve as a leading choice specifically for retrieval-focused applications like RAG

Using BGE Models via Sentence Transformers

Note the normalize_embeddings=True argument — BGE models are
specifically designed and evaluated with normalized (unit-length)
embeddings, which, as covered in the Euclidean Distance topic,
makes cosine similarity and Euclidean distance produce
equivalent rankings — following this convention is important
for getting the expected, benchmarked performance from the model.

BGE Model Size Variants

VariantGeneral Characteristics
bge-smallSmallest, fastest, lowest resource requirements
bge-baseBalanced tradeoff between speed and embedding quality
bge-largeHighest quality among the standard sizes, higher resource cost
This mirrors the same size-vs-performance tradeoff discussed
throughout this curriculum (Parameters and Model Size,
Scaling Laws) — larger BGE variants generally perform better
on benchmarks but require more compute and memory, and produce
slower inference, making the right choice dependent on your
specific latency, cost, and quality requirements.

Instruction-Aware Retrieval: A BGE Specialty

A distinctive characteristic of several BGE models is their
use of instruction prefixes specifically for queries in
retrieval tasks — the document/passage side is typically
embedded without any special prefix, but prepending a specific
instruction string to the QUERY has been shown to meaningfully
improve retrieval accuracy for these particular models. Always
check the specific model's documentation for its exact
recommended instruction format, since these details can vary
between BGE model versions.

BGE-M3: Multi-Functionality in One Model

BGE-M3 is a notable variant designed to support multiple
embedding capabilities within a single model:

- Multi-Linguality: strong support across 100+ languages
- Multi-Granularity: handles inputs ranging from short
  sentences to long documents effectively
- Multi-Functionality: can produce dense embeddings (the
  standard vector approach covered throughout this section),
  as well as sparse and multi-vector representations for
  more specialized retrieval techniques

This makes BGE-M3 a versatile choice for applications needing
flexibility across languages, document lengths, or retrieval
strategies within one consistent model family.

BGE Models vs Sentence Transformers (General)

AspectBGE ModelsGeneral Sentence Transformers Models
OriginDeveloped specifically by BAAIA broad ecosystem including many different model providers
Benchmark PerformanceFrequently ranks near the top on MTEB leaderboardsVaries significantly by specific model chosen
Instruction PrefixesRecommended for queries in several BGE variantsVaries by model; not universal
Multilingual SupportStrong, especially with BGE-M3Varies by specific model chosen
Access MethodUsed through the sentence-transformers library, like most open modelsThe library itself; hosts many different model families, including BGE
It's worth clarifying: "Sentence Transformers" (the previous
topic) refers to both a LIBRARY and a family of models built
by its original creators, while BGE is a separate model family
built by BAAI that happens to be USED through that same library
— they aren't strictly competing categories, but rather a
library and one of the many strong model families it supports.

Key Properties of BGE Models

  • BGE models are open-source, self-hostable embedding models developed by BAAI.
  • They're accessed through the standard sentence-transformers library, using familiar conventions.
  • Several BGE models recommend normalizing embeddings and using instruction prefixes for queries specifically.
  • BGE-M3 offers multilingual, multi-granularity, and multi-functionality support within a single model.
  • BGE models frequently achieve strong, competitive results on standard embedding benchmarks like MTEB.

Where Are BGE Models Used?

FieldApplication
RAG-Based ApplicationsHigh-quality document retrieval for LLM context
Multilingual Search SystemsLeveraging BGE-M3's strong cross-language support
Enterprise Document SearchSelf-hosted, privacy-preserving semantic search
Academic and Research BenchmarkingFrequently used as a strong baseline in embedding research
Long-Document RetrievalBGE-M3's multi-granularity support for varying document lengths

Advantages

  • Open-source and self-hostable, avoiding per-request API costs and data privacy concerns
  • Consistently strong performance on standard embedding benchmarks
  • Multiple size variants allow tuning the quality-vs-speed tradeoff for a given application
  • BGE-M3 provides notable flexibility across languages, document lengths, and retrieval techniques
  • Compatible with the widely-used and well-documented sentence-transformers library

Limitations

  • Requires self-hosting infrastructure, unlike a fully managed API-based option
  • Instruction prefix conventions add a small but important implementation detail to get right
  • Larger BGE variants require meaningful compute resources for good performance
  • Benchmark leaderboard rankings can shift over time as new models are released
  • Requires validating that a specific BGE variant performs well on your particular domain, not just general benchmarks

Real-World Examples

ApplicationBGE Model Use
Self-Hosted RAG SystemsUsing bge-base or bge-large for high-quality document retrieval
Multilingual Customer SupportUsing BGE-M3 to support search across multiple languages
Enterprise Knowledge BasesSelf-hosted BGE embeddings for internal, privacy-sensitive document search
Academic Research BenchmarkingComparing new embedding techniques against BGE as a strong baseline
Long-Document Search ApplicationsBGE-M3's multi-granularity handling of lengthy documents

Best Practices

  • Use normalize_embeddings=True when encoding with BGE models, following their designed convention.
  • Check the specific BGE model's documentation for its recommended query instruction prefix, if applicable.
  • Choose a BGE size variant based on your specific latency, cost, and quality requirements.
  • Consider BGE-M3 specifically for multilingual or variable-length document retrieval needs.
  • Benchmark BGE models against your own domain-specific data, not just general leaderboard rankings.

Interview Tip

A common interview question is:

"What distinguishes BGE models from other embedding models, and why might you use an instruction prefix only for queries and not for documents?"

A strong answer is:

BGE models, developed by BAAI, are open-source embedding models that consistently rank near the top of benchmarks like MTEB, and several variants recommend using an instruction prefix specifically for query embeddings during retrieval tasks — while document/passage embeddings are typically left as-is. This asymmetric approach helps the model better distinguish the specific ROLE each piece of text plays in a retrieval task: a query is a request looking for relevant information, while a document is the content being searched, and providing this signal specifically on the query side has been shown to improve retrieval accuracy for these models without needing to modify the potentially much larger set of documents being indexed.

Explaining the reasoning behind the query-only instruction convention makes your answer stronger.

Conclusion

BGE models offer a strong, open-source, self-hostable option in the embedding landscape, consistently performing well on standard benchmarks while integrating seamlessly with the sentence-transformers library already covered in this section. With BGE now covered, the next topic explores E5 Models, another leading open-source embedding family with its own distinctive conventions and strengths.