Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture

About

Generative large language models (LLMs) have become crucial for modern NLP research and applications across various languages. However, the development of foundational models specifically tailored to the Russian language has been limited, primarily due to the significant computational resources required. This paper introduces the GigaChat family of Russian LLMs, available in various sizes, including base models and instruction-tuned versions. We provide a detailed report on the model architecture, pre-training process, and experiments to guide design choices. In addition, we evaluate their performance on Russian and English benchmarks and compare GigaChat with multilingual analogs. The paper presents a system demonstration of the top-performing models accessible via an API, a Telegram bot, and a Web interface. Furthermore, we have released three open GigaChat models in open-source (https://huggingface.co/ai-sage), aiming to expand NLP research opportunities and support the development of industrial solutions for the Russian language.

GigaChat team: Mamedov Valentin, Evgenii Kosarev, Gregory Leleytner, Ilya Shchuckin, Valeriy Berezovskiy, Daniil Smirnov, Dmitry Kozlov, Sergei Averkiev, Lukyanenko Ivan, Aleksandr Proshunin, Ainur Israfilova, Ivan Baskov, Artem Chervyakov, Emil Shakirov, Mikhail Kolesov, Daria Khomich, Darya Latortseva, Sergei Porkhun, Yury Fedorov, Oleg Kutuzov, Polina Kudriavtseva, Sofiia Soldatova, Kolodin Egor, Stanislav Pyatkin, Dzmitry Menshykh, Grafov Sergei, Eldar Damirov, Karlov Vladimir, Ruslan Gaitukiev, Arkadiy Shatenov, Alena Fenogenova, Nikita Savushkin, Fedor Minkin• 2025

Related benchmarks

TaskDatasetResultRank
Coding ReasoningruLCB
Accuracy0.272
11
Advanced ReasoningT-Math
Accuracy14.2
11
Advanced ReasoningruAIME 2024
Accuracy10.2
11
Advanced ReasoningruMATH-500
Accuracy70.2
11
Advanced ReasoningruGPQA Diamond
Accuracy0.475
11
Advanced ReasoningruAIME 2025
Accuracy6.2
11
Advanced ReasoningVikhr Math
Accuracy37.2
11
Advanced ReasoningVikhr Physics
Accuracy24.5
11
TokenizationWikipedia Cyrillic-rich subsets (test)
Russian (ru)2.49
4
Showing 9 of 9 rows

Other info

Follow for update