IB-GRPO: Aligning LLM-based Learning Path Recommendation with Educational Objectives via Indicator-Based Group Relative Policy Optimization

About

Learning Path Recommendation (LPR) aims to generate personalized sequences of learning items that maximize long-term learning effect while respecting pedagogical principles and operational constraints. Although large language models (LLMs) offer rich semantic understanding for free-form recommendation, applying them to long-horizon LPR is challenging due to (i) misalignment with pedagogical objectives such as the Zone of Proximal Development (ZPD) under sparse, delayed feedback, (ii) scarce and costly expert demonstrations, and (iii) multi-objective interactions among learning effect, difficulty scheduling, length controllability, and trajectory diversity. To address these issues, we propose IB-GRPO (Indicator-Based Group Relative Policy Optimization), an indicator-guided alignment approach for LLM-based LPR. To mitigate data scarcity, we construct hybrid expert demonstrations via Genetic Algorithm search and teacher RL agents and warm-start the LLM with supervised fine-tuning. Building on this warm-start, we design a within-session ZPD alignment score for difficulty scheduling. IB-GRPO then uses the $I_{\epsilon+}$ dominance indicator to compute group-relative advantages over multiple objectives, avoiding manual scalarization and improving Pareto trade-offs. Experiments on ASSIST09 and Junyi using the KES simulator with a Qwen2.5-7B backbone show consistent improvements over representative RL and LLM baselines.

Shuai Wang, Yaoming Yang, Bingdong Li, Hao Hao, Aimin Zhou• 2026

Related benchmarks

Task	Dataset	Result
Learning Path Recommendation	Junyi L=20 (test)	Learning Effect (Ep)0.7743	9
Learning Path Recommendation	Junyi L=10 (test)	Learning Effect (Ep)0.6435	9
Learning Path Recommendation	Junyi L=5 (test)	Learning Effect (Ep)0.6115	9
Learning Path Recommendation	ASSIST09 L=20 (test)	Learning Effect (Ep)0.5911	9
Learning Path Recommendation	ASSIST09 L=10 (test)	Learning Effect (Ep)0.5139	9
Learning Path Recommendation	ASSIST09 L=5 (test)	Learning Effect (Ep)0.4222	9

Showing 6 of 6 rows

Other info

Follow for update

@wizwand_team Discord