Gradient-Guided Reward Optimization for Inference-time Alignment
About
Ensuring the reliability of Large Language Models (LLMs) under distribution drift requires inference-time adaptation. While inference-time alignment methods such as Best-of-$N$ and rejection sampling are widely used, they frame the task as a sampling-intensive, reward-guided search, leading to two key limitations: their performance is bounded by the base model's generation quality, and their reliance on imperfect reward models makes them vulnerable to reward hacking. To address these challenges, we introduce Gradient-Guided Reward Optimization (GGRO), a lightweight inference-time method that performs targeted, minimal intervention during decoding via gradient guidance. Specifically, GGRO monitors token-level entropy to identify high-uncertainty regions indicative of drift or misalignment. Upon detection, it responds by injecting nudging tokens, generated using gradient signals from an off-the-shelf reward model, to steer the generation trajectory rather than merely re-ranking samples. Experiments show that GGRO consistently improves inference-time alignment across safety, helpfulness, and reasoning benchmarks. It also increases coverage of high-quality responses and robustness to reward hacking, with minimal computational overhead. Code is available at https://github.com/lhk2004/GGRO.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Reasoning | MMLU-Pro | Accuracy54 | 264 | |
| Safety Evaluation | HEX-PHI | Attack Success Rate (ASR)26.2 | 107 | |
| Reasoning | ARC Challenge | Accuracy94.3 | 44 | |
| Safety Evaluation | XSTest | FRR1.2 | 35 | |
| Helpfulness | HH-RLHF | Gemini Score8.75 | 16 | |
| Pairwise Helpfulness Evaluation | HH-RLHF | Win Rate64.7 | 14 | |
| Adversarial Safety Evaluation | HEx-PHI (first 100 harmful prompts) | ASR22 | 8 |