Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents

About

Reinforcement learning with verifiable rewards improves reasoning and tool use, yet long-horizon language agents still learn unsupported evidence chains, belief drift, and shortcut actions that satisfy terminal checks. Existing process rewards are mostly correlational: they reward retrieval-, reflection-, or verification-like steps without estimating whether the step contributes to final verified success under a specified intervention. We propose CVT-RL, a constrained policy-gradient algorithm with dense verifiable rewards, intervention-validity gating, and a policy-conditioned counterfactual contribution (PCCC) estimator. Deletion, semantic substitution, evidence substitution, and tool-output perturbation define separate controlled interventions; continuations are sampled from a frozen reference policy, and a selection-adjusted doubly robust estimator augments the advantage. Belief control uses only prefix-observable labels, while an augmented Lagrangian constrains unsupported claims, skipped verification, tool tampering, and unsafe calls. On long-context QA, ALFWorld, ScienceWorld, and web/tool tasks, CVT-RL improves average task success from 71.8% for compute-matched non-causal RL and 75.4% for an information-matched counterfactual-process baseline to 78.9%, improves evidence F1 from 78.9 to 82.8 over the information-matched baseline, and reduces measured hacking from 7.2% to 3.9%. Independent human audit estimates 4.6% hacking for CVT-RL versus 8.1% for the information-matched baseline, and adaptive detector-evasion attacks raise hacking only to 7.1%. Stratified bootstrap and mixed-effects tests give p<0.01 after Holm correction for all primary metrics. Carefully scoped counterfactual credit, paired with validity gating, diagnostics, and verifiable constraints, provides a reproducible route toward more reliable long-horizon RL for language agents.

Renwei Meng• 2026

Related benchmarks

TaskDatasetResultRank
Interactive Decision-makingAlfWorld
Overall Success Rate77.4
398
Scientific ReasoningScienceWorld
Success Rate75.9
26
Agentic Performance AggregateCombined LCtx, ALFWorld, ScienceWorld, Web/Tool
Mean Task Success78.9
10
Question AnsweringLong-context QA
Task Success82.7
10
Web Navigation and Tool UseWeb Tool
Task Success79.6
10
Embodied text agentsALFWorld 134 (test)
Success Rate77.4
3
Long-context QARULER 1200 (test)
Success Rate84.1
3
Long-context QALongBench 1000 (test)
Success Rate81.3
3
Long-context QALooGLE 900 (test)
Success Rate82.8
3
Scientific interactionScienceWorld 1400 (test)
Success Rate75.9
3
Showing 10 of 14 rows

Other info

Follow for update