Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs

About

Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address implicit, domain-specific risks. To investigate this gap, we introduce a dataset of 3,000 annotated queries spanning education, finance, and management. Evaluations across 14 leading LLMs reveal a concerning vulnerability: an average jailbreak success rate of 57.8\%. In response, we propose MENTOR, a metacognition-driven self-evolution framework. MENTOR performs metacognitive self-assessment, using strategies such as perspective-taking and consequential reasoning to uncover latent model misalignments. The resulting reflections are distilled into dynamic rule-based knowledge graphs, from which retrieved rules are converted into activation-level steering signals to guide internal representations during inference. Experiments demonstrate that MENTOR substantially reduces attack success rates across all tested domains and outperforms existing safety alignment methods. The code and dataset for MENTOR are available at: https://anonymous.4open.science/r/MENTOR-Evo.

Liang Shan, Kaicheng Shen, Wen Wu, Zhenyu Ying, Chaochao Lu, Yan Teng, Jingqi Huang, Qingshan Liu, Guangze Ye, Guoqing Wang, Jie Zhou, Liang He• 2025

Related benchmarks

TaskDatasetResultRank
Safety and Utility EvaluationFinance
JSR0.004
44
Safety and Utility EvaluationManagement subset
JSR Score0.78
44
Safety and Utility EvaluationEducation subset
JSR65.6
44
Jailbreak defense and Utility evaluationImplicit risk dataset Education
Jailbreak Success Rate (JSR)6.4
30
Jailbreak defense and Utility evaluationImplicit risk dataset Finance
Jailbreak Success Rate (JSR)0.4
30
Jailbreak defense and Utility evaluationImplicit risk dataset Management
Jailbreak Success Rate (JSR)0.8
30
Human EvaluationEducation subset
Win Rate85
3
Human EvaluationFinance
Win Rate97
3
Human EvaluationManagement subset
Win Rate93
3
Jailbreak Success RateAdvBench--
3
Showing 10 of 15 rows

Other info

Follow for update