Home
Softono

Bge M3 Onnx

Open source Apache-2.0 Jupyter Notebook
40
Stars
5
Forks
8
Issues
0
Watchers
3 months
Last Commit

 About Bge M3 Onnx

bge-m3-onnx is an ONNX implementation of the BGE-M3 multilingual embedding model and tokenizer, enabling native execution in C, Java, and Python. This solution generates all three specific embedding types supported by BGE-M3: dense, sparse, and ColBERT vectors. It is designed for developers requiring high-performance, multi-vector retrieval capabilities without external dependencies or internet connectivity. The package supports cross-platform deployment and includes CUDA GPU acceleration for reduced latency in local embedding generation. Key features include full control over the embedding pipeline, offline operation, and verified cross-language compatibility ensuring consistent vector outputs across .NET, Java, and Python environments. The repository provides complete tooling for converting the original FlagEmbedding BAAI/bge-m3 model to ONNX format, including a Jupyter notebook guide, reference embedding generators, and test scripts to validate consistency across languages. Samples are provided for each su

Platforms

Web Self-hosted

Languages

Jupyter Notebook

Links

Need Help Installing Bge M3 Onnx?

We provide expert installation service for this software. Our team will install, configure, and secure Bge M3 Onnx on your server. plans start at just $30.

Bge M3 Onnx

View on GitHub

BGE-M3 ONNX

Build Status HuggingFace Model Build

This repository demonstrates how to convert the complete BGE-M3 model to ONNX format and use it in multiple programming languages with full multi-vector functionality.

image

Key Features

  • Generate all three BGE-M3 embedding types: dense, sparse, and ColBERT vectors
  • Reduced latency with local embedding generation
  • Full control over the embedding pipeline with no external dependencies
  • Works offline without internet connectivity requirements
  • Cross-platform compatibility (C#, Java, Python)
  • CUDA GPU acceleration support

Repository Structure

  • bge-m3-to-onnx.ipynb - Jupyter notebook demonstrating the BGE-M3 conversion process
  • /samples/dotnet - C# implementation
  • /samples/java - Java implementation
  • /samples/python - Python implementation
  • generate_reference_embeddings.py - Script to generate reference embeddings for cross-language testing
  • run_tests.sh and run_tests.ps1 - Test scripts for Linux/macOS and Windows

Getting Started

  1. Clone this repository:

    git clone https://github.com/yuniko-software/bge-m3-onnx.git
    cd bge-m3-onnx
    
  2. Get the BGE-M3 ONNX models:

    • Option 1: Download from releases (recommended)

      • Check the repository releases and download onnx.zip
      • It already contains the bge-m3 embedding model and its tokenizer
    • Option 2: Generate yourself using the notebook

      • Open and run bge-m3-to-onnx.ipynb - this is the most important file in the repository
      • The notebook demonstrates how to convert BGE-M3 from FlagEmbedding to ONNX format
      • This will create bge_m3_tokenizer.onnx, bge_m3_model.onnx, and bge_m3_model.onnx_data in the /onnx folder

    Note: This repository uses BAAI/bge-m3 as the embedding model with its XLM-RoBERTa tokenizer.

  3. Generate reference embeddings (optional):

    • Run python generate_reference_embeddings.py to create reference embeddings for testing

    Python dependencies are managed in requirements.txt:

    pip install -r requirements.txt
    
  4. Run the samples:

    • Once you have the ONNX models in the /onnx folder, you can run any sample
    • Try the .NET sample in /samples/dotnet or the Java sample in /samples/java
  5. Verify cross-language embeddings (optional):

    • To ensure that .NET and Java embeddings match the Python-generated embeddings, you can run:

    • On Linux/macOS:

      chmod +x run_tests.sh
      ./run_tests.sh
      
    • On Windows:

      ./run_tests.ps1
      

    Note: These scripts require Python, .NET, Java, and Maven to be installed.

CUDA Support

This BGE-M3 ONNX model supports CUDA GPU acceleration for improved performance. To enable CUDA support:

Python

Install the ONNX Runtime with CUDA support:

Resource: ONNX Runtime CUDA Execution Provider Requirements

This model is compatible with:

C# and Java

For C# and Java implementations, you need to install CUDA and cuDNN separately:

CUDA Installation:

cuDNN Installation:

Python Example

from bge_m3_embedder import create_cpu_embedder, create_cuda_embedder

# Create CPU-optimized embedder
embedder = create_cpu_embedder("onnx/bge_m3_tokenizer.onnx", "onnx/bge_m3_model.onnx")

# Generate all three embedding types
result = embedder.encode("Hello world!")

print(f"Dense: {len(result['dense_vecs'])} dimensions")
print(f"Sparse: {len(result['lexical_weights'])} tokens")  
print(f"ColBERT: {len(result['colbert_vecs'])} vectors")

# Clean up resources
embedder.close()

# For CUDA acceleration
cuda_embedder = create_cuda_embedder("onnx/bge_m3_tokenizer.onnx", "onnx/bge_m3_model.onnx", device_id=0)
result = cuda_embedder.encode("Hello world!")
cuda_embedder.close()

# See full implementation in samples/python

C# Example

using BgeM3.Onnx;

// Create CPU-optimized embedder
using var embedder = M3EmbedderFactory.CreateCpuOptimized(tokenizerPath, modelPath);

// Generate all embedding types
var result = embedder.GenerateEmbeddings("Hello world!");

Console.WriteLine($"Dense: {result.DenseEmbedding.Length} dimensions");
Console.WriteLine($"Sparse: {result.SparseWeights.Count} tokens");
Console.WriteLine($"ColBERT: {result.ColBertVectors.Length} vectors");

// For CUDA acceleration
using var cudaEmbedder = M3EmbedderFactory.CreateCudaOptimized(tokenizerPath, modelPath, deviceId: 0);
var cudaResult = cudaEmbedder.GenerateEmbeddings("Hello world!");

// See full implementation in samples/dotnet

Java Example

import com.yunikosoftware.bgem3onnx.*;

// Create CPU-optimized embedder
try (M3Embedder embedder = M3EmbedderFactory.createCpuOptimized(tokenizerPath, modelPath)) {
    // Generate all embedding types
    M3EmbeddingOutput result = embedder.generateEmbeddings("Hello world!");
    
    System.out.println("Dense: " + result.getDenseEmbedding().length + " dimensions");
    System.out.println("Sparse: " + result.getSparseWeights().size() + " tokens");
    System.out.println("ColBERT: " + result.getColBertVectors().length + " vectors");
}

// For CUDA acceleration
try (M3Embedder cudaEmbedder = M3EmbedderFactory.createCudaOptimized(tokenizerPath, modelPath, 0)) {
    M3EmbeddingOutput result = cudaEmbedder.generateEmbeddings("Hello world!");
    // Process CUDA results
}

// See full implementation in samples/java

If you find this project useful, please consider giving it a star on GitHub!

Your support helps make this project more visible to other developers who might benefit from BGE-M3's complete multi-vector functionality.