AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation
About
The generation of anchor-style product promotion videos presents promising opportunities in e-commerce, advertising, and consumer engagement. Despite advancements in pose-guided human video generation, creating product promotion videos remains challenging. In addressing this challenge, we identify the integration of human-object interactions (HOI) into pose-guided human video generation as a core issue. To this end, we introduce AnchorCrafter, a novel diffusion-based system designed to generate 2D videos featuring a target human and a customized object, achieving high visual fidelity and controllable interactions. Specifically, we propose two key innovations: the HOI-appearance perception, which enhances object appearance recognition from arbitrary multi-view perspectives and disentangles object and human appearance, and the HOI-motion injection, which enables complex human-object interactions by overcoming challenges in object trajectory conditioning and inter-occlusion management. Extensive experiments show that our system improves object appearance preservation by 7.5\% and doubles the object localization accuracy compared to existing state-of-the-art approaches. It also outperforms existing approaches in maintaining human motion consistency and high-quality video generation. Project page including data, code, and Huggingface demo: https://github.com/cangcz/AnchorCrafter.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Video Generation | User Study | Interaction Plausibility Score6.55 | 16 | |
| Human-Object Interaction Video Generation | AnchorCrafter (test) | MS (Motion Score)99.23 | 14 | |
| Human-Object Interaction Video Generation | Mani4D (test) | Obj-IoU64.61 | 7 | |
| HOI Video Generation | HOI video generation (test) | AES Score44.8 | 7 | |
| HOI Video Generation | SparseHOI-5K | FID258.4 | 6 | |
| HOI Video Generation | AnchorCrafter | FID91.9 | 4 | |
| User Study | SparseHOI-5K 1.0 (full) | Videos Count356 | 3 | |
| Text+Reference+Pose-to-Video (RP2V) Generation | HOIVG-Bench 1.0 (test) | TA2.669 | 3 | |
| Human-Object Interaction Video Generation | Mani4D (Novel Human Reference) | Obj IoU24.35 | 2 |