| Task Name | Dataset Name | SOTA Result | Trend | |
|---|---|---|---|---|
| Workflow Orchestration | Video | Success Rate93 | 36 | |
| Membership Inference | Video modality | AUC-ROC100 | 16 | |
| Video Analytics | Video | Cost per 1K Requests0.49 | 15 | |
| Video Super-Resolution | video 33-frame 720x1280 | Inference Time (s)4.97 | 13 | |
| Neural Video Representation | Video per-frame | GFLOPs64.92 | 12 | |
| Video reasoning | Video-R1 | VSI44.3 | 12 | |
| Face Video Restoration | 161-frame 512x512 video | FPS1.988 | 10 | |
| Motion Tracking | Video Unseen | Total Performance82.9 | 8 | |
| Video Super-Resolution | 30-frame 2K Video (test) | Inference Time (min)0.77 | 8 | |
| Sequential Recommendation | Video (test) | NDCG@106.436 | 8 | |
| Video Super-Resolution | video 1920x1080 (21-frame sequence) | Step Count50 | 8 | |
| Sequential Recommendation | Video | NDCG@52.17 | 8 | |
| Future item recommendation | Video | Recall11.3 | 7 | |
| Recommendation | Video | Hit Rate@1054.3 | 6 | |
| Video Semantic Segmentation | 1024 x 512 resolution (video) | Speed (FPS)18.15 | 6 | |
| Text-to-Video | Video 480x832x97 | Inference Time (s)132.23 | 5 | |
| Point Tracking | 24-frame video | Throughput23,405.71 | 5 | |
| Session-based recommendation | VIDEO | Recall@2066.24 | 5 | |
| Instruction-based Video Editing | Video 480x832x97 | Inference Time (s)166.3 | 4 | |
| User Cold-start Recommendation | Video | Recall@209.22 | 4 | |
| Visual Dubbing | Video 3-second 25fps 512x512 resolution | Inference Time (s)1 | 4 | |
| Object Detection | video (train) | Accuracy93 | 4 | |
| Recommendation | Video | Processing Time (sec)10 | 3 | |
| Misalignment reduction | Video #5 | ITF (dB)21.66 | 3 | |
| Misalignment reduction | Video #4 | ITF (dB)19.26 | 3 |