Bielik 11B v3: Multilingual Large Language Model for European Languages
About
We present Bielik 11B v3, a state-of-the-art language model highly optimized for the Polish language, while also maintaining strong capabilities in other European languages. This model extends the Mistral 7B v0.2 architecture, scaled to 11B parameters via depth up-scaling. Its development involved a comprehensive four-stage training pipeline: continuous pre-training, supervised fine-tuning (SFT), Direct Preference Optimization (DPO), and reinforcement learning. Comprehensive evaluations demonstrate that Bielik 11B v3 achieves exceptional performance. It significantly surpasses other specialized Polish language models and outperforms many larger models (with 2-6 times more parameters) on a wide range of tasks, from basic linguistic understanding to complex reasoning. The model's parameter efficiency, combined with extensive quantization options, allows for effective deployment across diverse hardware configurations. Bielik 11B v3 not only advances AI capabilities for the Polish language but also establishes a new benchmark for developing resource-efficient, high-performance models for less-represented languages.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Emotional Intelligence | Polish EQ-Bench | Overall Score71.2 | 106 | |
| Polish Text Understanding | CPTUB | Overall Avg3.73 | 98 | |
| Linguistic and Cultural Competency | Polish Linguistic and Cultural Competency Benchmark (PLCC) | Avg Score71.83 | 52 | |
| Multilingual Language Proficiency | INCLUDE base 44 | Average Score64.8 | 46 | |
| Polish Instruction Following | Open PL LLM Leaderboard | Average Score65.93 | 45 | |
| Large Language Model Evaluation | Open PL LLM Leaderboard instruction-tuned | Overall Average Score65.93 | 44 | |
| Large Language Model Evaluation | Open LLM Leaderboard | Average Score72.45 | 41 | |
| Reading Comprehension | Belebele 28 European languages | Overall Score82.98 | 34 | |
| Medical Question Answering | Polish Board Certification Examinations | Average Score50.21 | 30 | |
| Linguistic Implicatures Decoding | Open PL LLM Leaderboard Implicatures component base models | Average Score55.16 | 30 |