Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SAGE: Retain-Aware Post-Hoc Sanitization of Final Unlearning Vector

About

Large Language Model (LLM) unlearning aims to remove undesirable knowledge or behaviors while preserving retained capabilities. Current unlearning methods all involve a trade-off between unlearning and retention. We have found that the retention activation bias can also be used to quantify the damage an unlearning method inflicts on retention, without considering the specific implementation of the unlearning process. This allows us to restore retention performance for any unlearning method using a post-hoc approach. Therefore, we propose a complementary post-hoc setting to sanitize the final update vector without rerunning the original unlearning pipeline. In this setting, we design SAGE, Spectral Activation-GEometry Sanitization, a source-agnostic correction for final unlearning updates. SAGE collects real module inputs from a small retain proxy, extracts their dominant activation geometry, and solves a source-anchored optimization objective in closed form, which suppresses update components aligned with high-energy retained directions while preserving the source method's forgetting carrier. Across multiple unlearning methods, model scales, and benchmarks, SAGE consistently relieves the retain-forget trade-off, identifying post-hoc sanitization of final vectors as a practical and underexplored axis for machine unlearning.

Jingyuan Zhang, Yucheng Bai, Peixi Wen, Zhehao Huang, Zhengbao He, Hanling Tian, Xinwen Cheng, Haiyin Ran, Xiaolin Huang• 2026

Related benchmarks

TaskDatasetResultRank
Machine UnlearningTOFU 10% forget
Privacy Leakage14.05
72
Machine UnlearningMUSE NEWS
ES Unlearning0.011
14
Machine UnlearningMUSE Books
ES Un.10.6
14
Machine UnlearningWMDP cyber
Unlearning Accuracy42.6
13
Showing 4 of 4 rows

Other info

Follow for update