Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents

About

We present LuckyStar 111B, a 111B-parameter hybrid reasoning model developed through a collaboration between Cohere and LG CNS for Korean-English enterprise agents under practical memory and serving constraints. The model trains from Cohere's fully post-trained Command A model rather than a new pretraining run, and uses preamble conditioning to switch between concise non-reasoning behavior and longer tool-oriented reasoning. We study four choices for scaling tool-using agents efficiently: multilingual supervised fine-tuning, reinforcement learning with verifiable rewards for multi-step tool-use tasks, language-consistency rewards for Korean user-facing responses, and 4-bit quantization for single-GPU serving. The adapted model improves mathematical reasoning, function calling, and agentic natural-language-to-SQL (NL2SQL) performance while preserving general Korean and English instruction-following quality. These results provide a practical recipe and failure-mode analysis for adapting post-trained multilingual models to verifiable agentic workflows under memory-constrained deployment.

Utsav Garg, Sungjin Hong, Jason Jung, Justin Lee, Shaan Desai, Joon Hee Kim, Anirudh Shrinivason, Edmond Wen, Susie Park• 2026

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningMATH 500
Pass@196
12
Instruction FollowingIFEval Korean
Accuracy78.3
6
Knowledge EvaluationKMMLU (test)
Accuracy68.6
6
Mathematical ReasoningMATH 500 (Korean)
Pass@1 Accuracy95.6
6
Instruction FollowingIFEval English
Accuracy89.4
6
Mathematical ReasoningAIME Korean 2024
Pass@1 Accuracy69.3
6
Mathematical ReasoningAIME English 2024
Pass@1 Accuracy73.7
6
Science ReasoningARC-C English (test)
Accuracy93.8
6
Conversation QualityMT-Bench English
Score8.5
6
Science ReasoningARC-C Korean (test)
Accuracy89.2
6
Showing 10 of 14 rows

Other info

Follow for update