Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling

About

Reliable reward and preference signals are critical for evaluating and optimizing large language models on open-ended tasks. Rubric-based judges offer a transparent way to decompose such judgments into explicit evaluation criteria, but existing annotation-free rubric generators typically rely on a single generic evaluator. As a result, they may overlook important dimensions of human preference, a failure mode we term dimensional blind spots. To address this limitation, we propose Multi-Role Rubric Generation (MRRG), a training-free and reference-free framework that elicits evaluation criteria from multiple complementary roles and consolidates them into an auditable rubric-based scorer. This scorer can be used both to validate pairwise preferences and to provide rewards for GRPO-style Reinforcement Learning with Verifiable Rewards (RLVR). Experiments on preference validation benchmarks show that MRRG consistently outperforms single-role rubric generation baselines across multiple backbone models. Further RLVR experiments demonstrate that MRRG yields a stronger reward signal for improving open-ended generation.

Dazhi Fu, Jiuding Yang, Yiwen Guo, Jicong Fan• 2026

Related benchmarks

TaskDatasetResultRank
Preference PredictionJudgeBench
Positional Consistent Accuracy74.8
30
Preference ValidationRewardBench 2
Accuracy74.5
20
Preference ValidationPPE
Accuracy57.8
20
Post-RL EvaluationBiGGen-Bench
Accuracy63.7
5
Post-RL EvaluationHealthBench Hard
Accuracy32.1
5
Showing 5 of 5 rows

Other info

Follow for update