Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback
About
We study multi-domain LLM training in which two models, each stronger in a different domain, co-evolve by tutoring each other through on-policy feedback. Unlike one-way distillation or single-model fine-tuning, our goal is mutual Pareto improvement: each model improves across domains without losing its original strength. To this end, we propose On-Policy Co-Distillation (OPCoD), where each student's self-distillation is conditioned on its own correct rollout and feedback from its peer. To make feedback exchange effective, OPCoD uses cognizance-based gating to decide when to give feedback and feedback anchoring to ground feedback in the problem. On Science Q\&A tasks, OPCoD consistently outperforms baselines and achieves Pareto improvement across all evaluated domain pairs and students.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Multi-domain Science Q&A | SciKnowEval Mat-Phys (test) | Accuracy (Mat. Sci.)70.5 | 8 | |
| Multi-domain Science Q&A | SciKnowEval Chem-Mat (test) | Chemistry Accuracy71.3 | 8 | |
| Multi-domain Science Q&A | SciKnowEval Phys-Chem (test) | Accuracy (Physics)58.8 | 8 |