Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Brain-Inspired Stochastic Joint Embedding Representation Learning

About

Representation learning is one of the key research topics in machine learning, and the framework of self-supervised learning (SSL) has revolutionized computer vision. However, these approaches have not yet fully leveraged insights from biological visual processing systems. In this paper, we introduce PhiNet v2, a novel architecture that processes temporal visual input (i.e., sequences of images) without relying on strong data augmentation, enabling it to learn robust visual representations in a manner similar to human visual processing. Our learning objective is derived from variational inference. Through extensive experiments, we demonstrate that PhiNet v2 achieves competitive performance compared to state-of-the-art vision representation models, including RSP and CropMAE, while retaining the ability to learn effectively from sequential input without strong data augmentation. This work represents a step toward more biologically plausible computer vision systems that process visual information in a manner more aligned with human cognitive processes.

Makoto Yamada, Kian Ming A. Chai, Ayoub Rhim, Satoki Ishikawa, Mohammad Sabokrou, Yao-Hung Hubert Tsai• 2025

Related benchmarks

TaskDatasetResultRank
Video segmentationDAVIS
J&F Score60.1
53
Video Part SegmentationVIP
mIoU0.331
48
Pose TrackingJHMDB
PCK@0.145
20
Showing 3 of 3 rows

Other info

Follow for update