SEA-Embedding: Open and Reproducible Text Embeddings for Southeast Asia
About
Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP. However, most recent state-of-the-art embedding models are not reproducible because they rely on closed or undisclosed training data, and they remain insufficiently robust for Southeast Asian languages. We present SEA-Embedding, a fully open and reproducible text-embedding pipeline for Southeast Asian languages trained only on publicly available data, and use it to study three core factors of robust embedding design: data composition, training objective, and base encoder initialization. SEA-Embedding achieves state-of-the-art results on SEA-BED while enabling systematic and reproducible analysis of robust text embeddings for the region.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Text Embedding | MTEB | Classification Score77.8 | 82 | |
| Text Embedding | SEA-BED | Accuracy (Indonesian)80.8 | 11 | |
| Text Embedding Evaluation | CMTEB | Classification Score71.3 | 9 | |
| Text Embedding | SEA-BED | Lang-Avg0.8 | 6 |