Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Korean Culture into LLM Alignment: Toward Cultural Coherence

About

Cultural-aspect work on large language models is dominated by a negative target: which outputs to suppress. We argue that a constructive counterpart is also needed, a working definition of what a culturally coherent response is rather than only what it must avoid, and instantiate it for Korean. We design an alignment-data pipeline around a prompt-based LLM seed generator that expands a Korean harm taxonomy, with a Korean-culturally-adapted safe-response policy at its centre: a per-category guideline grounded in Korean legal frameworks, social norms, and interpretive conventions, against which three frontier models each produce a candidate response. DPO fine-tuning on the resulting triplets improves the Korean cultural safe rate across six open-weight LLMs while causing no large degradation on Korean general-capability benchmarks, and qualitative outputs show fine-tuned models naming Korean statutes and institutional procedures and, where appropriate, supplying constructive Korean-context information alongside refusal.

MinJae Jung, Minwoo Kim• 2026

Related benchmarks

TaskDatasetResultRank
General Language UnderstandingKMMLU
Overall Score57.45
16
Code GenerationHumanEval+
Accuracy77.44
12
Mathematical ReasoningHRM8K
Accuracy (%)48.3
12
Multi-turn Instruction FollowingKo-MT-Bench
Overall Score (1-10)8.21
12
Safety and Cultural Bias EvaluationKoBBQ
Accuracy89.18
12
Safety EvaluationKorset
Safe Rate88.97
12
Showing 6 of 6 rows

Other info

Follow for update