Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Towards Healthy Evolution: Exploring the Role and Mechanisms of Human-Agent Interaction in Self-Evolving Systems

About

Self-evolving agents improve through continual self-play and self-generated learning signals, but autonomous evolution can also cause capability degradation and safety drift. Although human feedback has proven effective for static and post-trained agents, its role in self-evolving systems remains underexplored. We introduce Agent Norm Correction through Human-like Oversight and Review (ANCHOR), an LLM-based framework that simulates human supervision and delivers feedback at various phases of self-evolution. With ANCHOR, we evaluate two representative open-source self-evolving agent systems across coding, mathematical reasoning, and safety. Our results show that even limited supervision substantially mitigates safety degradation while preserving stable performance on core evolutionary objectives. Further analysis shows that supervision over the output verification phase is the most effective for intervention, whereas increasing supervision frequency yields diminishing returns. These findings provide empirical evidence and practical guidance for designing more stable, controllable, and human-aligned self-evolving agent systems.

Dianxing Shi, Bowen Wang, Junqi He, Junhao Chen, Yuta Nakashima• 2026

Related benchmarks

TaskDatasetResultRank
Agent Action PerformanceASR
ASR Average22.5
18
Code GenerationCode
Code Avg65.1
18
Reasoning / RewardRR
RR Average73.2
18
Human Safety / AlignmentHS
HS Average1.82
18
Showing 4 of 4 rows

Other info

Follow for update