Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Forget, Anticipate and Adapt: Test Time Training for Long Videos

About

Test Time Training (TTT) is a mechanism in which a model adapts to an incoming test-sample by performing some self-supervised (SSL) task and updating its weights even during inference. This procedure does not require labels at test-time. This paper focuses on TTT for long-videos. A major concern with existing approaches is: 1) they perform TTT updates using a sliding window containing frames in the past, whose compute increases linearly with the size of window. This becomes computationally intractable when the videos are hours long. 2) TTT is performed even when temporally close frames look similar, thereby consuming a lot of compute. We present the Frame Forgetting Network (FFN) that: 1) operates on only three frames within the sliding window, namely the frame that exits, the current frame and the frame after that. The model still manages to retain temporal context and work for hours long-videos; 2) mathematically define a surprise metric: how much new information the incoming frame contains with respect to the past seen frame. This facilitates determining how to modify the effective window size during TTT and constitutes the core mechanism of an adaptive windowing algorithm. Additionally, we curate a dataset EpicTours containing up to 3 hour long videos of walking city-tours, whereas earlier datasets on this problem were only 5 min long. We demonstrate FFNs empirical effectiveness on dense-segmentation, video classification tasks, generalization to depth-estimation, and multi-hour long videos.

Rajat Modi, Sebastian Noel, Xin Liang, Yogesh Singh Rawat• 2026

Related benchmarks

TaskDatasetResultRank
Action ClassificationUCF101
Top-1 Accuracy95
167
Video Depth EstimationBONN
AbsRel5.1
139
Monocular Video Depth EstimationNYU Video v2
Absolute Relative Error0.049
27
Video Depth EstimationScanNet
AbsRel6.2
15
Video Depth EstimationKITTI
AbsRel5.9
15
Video Depth EstimationSintel ~50 frames
AbsRel23
15
Semantic segmentationKITTI-STEP (val)
mIoU57.3
12
Semantic segmentationKITTI-STEP (test)
mIoU59.5
12
Instance SegmentationCOCO Videos
AP45.1
10
Panoptic SegmentationCOCO Videos
PQ29.6
10
Showing 10 of 14 rows

Other info

Follow for update