Resonant Minds: Closed-Loop Social Avatars with Theory of Mind
About
Creating lifelike digital humans with genuine social intelligence requires unifying cognitive reasoning and multimodal generation within a coherent framework. Current approaches treat these as separate tasks: Large Language Models excel at dialogue but lack embodied expression, while diffusion-based talking head models achieve visual fidelity but ignore social cognition. To bridge this gap, we propose a closed-loop dual-agent framework integrating perception, social reasoning, and expression into a continuous interaction cycle. The perception module analyzes partners' multimodal behaviors from video, while the social reasoning module infers hidden mental states through Theory of Mind and selects responses via an ensemble mechanism. The expression module then generates emotion-controllable videos that jointly synthesize speaker speech and facial expressions with listener reactive behaviors, capturing bidirectional dynamics absent in prior work. We further construct a hierarchical Persona-Scenario dataset with psychologically grounded personas and private social goals to support evaluation under information asymmetry. Experiments on this dataset demonstrate competitive or superior performance on both dialogue quality and video generation metrics. Notably, our method surpasses even the full-information Script mode on key dialogue quality dimensions, suggesting that explicit mental state inference under uncertainty can elicit more thoughtful dialogue than unrestricted information access. Project page: https://resonantminds.github.io/.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Audio Quality Assessment | URO-Bench | Open-D61 | 7 | |
| Talking head video generation | URO-Bench | Lip LMD22.39 | 7 | |
| Dual-agent Video Emotion Evaluation | URO-Bench | Emo Score36.04 | 5 | |
| Video Quality Evaluation | Video Quality User Study 6 scenario types, 8 emotion categories | Emotional Expression4.5 | 4 | |
| Dialogue Goal Completion | Persona-Scenario | Believability (Sotopia-Eval)9.19 | 3 | |
| Dialogue Naturalness | Persona-Scenario | GPT-Score: Fluency96.71 | 3 | |
| Dialogue Quality Evaluation | Persona-Scenario | Believability3.5 | 3 |