PERM: Psychology-grounded Empathetic Reward Modeling for Large Language Models

About

Large Language Models (LLMs) are increasingly deployed in human-centric applications, yet they often fail to provide substantive emotional support. While Reinforcement Learning (RL) has been utilized to enhance empathy of LLMs, existing reward models typically evaluate empathy from a single perspective, overlooking the inherently bidirectional interaction nature of empathy between the supporter and seeker as defined by Empathy Cycle theory. To address this limitation, we propose Psychology-grounded Empathetic Reward Modeling (PERM). PERM operationalizes empathy evaluation through a bidirectional decomposition: 1) Supporter perspective, assessing internal resonation and communicative expression; 2) Seeker perspective, evaluating emotional reception. Additionally, it incorporates a bystander perspective to monitor overall interaction quality. Extensive experiments on a widely-used emotional intelligence benchmark and an industrial daily conversation dataset demonstrate that PERM outperforms state-of-the-art baselines by over 10\%. Furthermore, a blinded user study reveals a 70\% preference for our approach, highlighting its efficacy in generating more empathetic responses. Our code, dataset, and models are available at https://github.com/ZhengWwwq/PERM.

Chengbing Wang, Wuqiang Zheng, Yang Zhang, Fengbin Zhu, Junyi Cheng, Yi Xie, Wenjie Wang, Fuli Feng• 2026

Related benchmarks

Task	Dataset	Result
Multi-turn role-play	ED	Success Rate (SR)91.8	12
Multi-turn role-play	MSD	Success Rate (SR)82.1	12
Multi-turn role-play	MedD	Success Rate (SR)94	12
Multi-turn role-play	ICLR	Success Rate (SR)88.3	12
Emotional Intelligence Evaluation	EQ-Bench3	Overall Score66.6	12

Showing 5 of 5 rows

Other info

Follow for update

@wizwand_team Discord