Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

Bielik 11B v3: Multilingual Large Language Model for European Languages

About

We present Bielik 11B v3, a state-of-the-art language model highly optimized for the Polish language, while also maintaining strong capabilities in other European languages. This model extends the Mistral 7B v0.2 architecture, scaled to 11B parameters via depth up-scaling. Its development involved a comprehensive four-stage training pipeline: continuous pre-training, supervised fine-tuning (SFT), Direct Preference Optimization (DPO), and reinforcement learning. Comprehensive evaluations demonstrate that Bielik 11B v3 achieves exceptional performance. It significantly surpasses other specialized Polish language models and outperforms many larger models (with 2-6 times more parameters) on a wide range of tasks, from basic linguistic understanding to complex reasoning. The model's parameter efficiency, combined with extensive quantization options, allows for effective deployment across diverse hardware configurations. Bielik 11B v3 not only advances AI capabilities for the Polish language but also establishes a new benchmark for developing resource-efficient, high-performance models for less-represented languages.

Krzysztof Ociepa, {\L}ukasz Flis, Remigiusz Kinas, Krzysztof Wr\'obel, Adrian Gwo\'zdziej• 2025

Related benchmarks

TaskDatasetResultRank
Linguistic and Cultural CompetencyPolish Linguistic and Cultural Competency Benchmark (PLCC)
Avg Score71.83
52
Large Language Model EvaluationOpen PL LLM Leaderboard instruction-tuned
Overall Average Score65.93
44
Emotional IntelligencePolish EQ-Bench
Overall Score71.2
34
Polish Text UnderstandingCPTUB
Overall Avg3.73
31
Linguistic Implicatures DecodingOpen PL LLM Leaderboard Implicatures component base models
Average Score55.16
30
Medical Knowledge PerformancePolish Board Certification Examinations (test)
Average Score (%)50.21
29
Language UnderstandingINCLUDE base 44
Average Score64.8
21
Large Language Model EvaluationOpen LLM Leaderboard
Average Score72.45
19
Large Language Model EvaluationOpen LLM Leaderboard v1 (test)
Average Score68.45
14
Reading ComprehensionBelebele 28 European languages
Overall Score82.98
10
Showing 10 of 11 rows

Other info

Follow for update