Topology-aware Convolutional Neural Network for Efficient Skeleton-based Action Recognition

About

In the context of skeleton-based action recognition, graph convolutional networks (GCNs) have been rapidly developed, whereas convolutional neural networks (CNNs) have received less attention. One reason is that CNNs are considered poor in modeling the irregular skeleton topology. To alleviate this limitation, we propose a pure CNN architecture named Topology-aware CNN (Ta-CNN) in this paper. In particular, we develop a novel cross-channel feature augmentation module, which is a combo of map-attend-group-map operations. By applying the module to the coordinate level and the joint level subsequently, the topology feature is effectively enhanced. Notably, we theoretically prove that graph convolution is a special case of normal convolution when the joint dimension is treated as channels. This confirms that the topology modeling power of GCNs can also be implemented by using a CNN. Moreover, we creatively design a SkeletonMix strategy which mixes two persons in a unique manner and further boosts the performance. Extensive experiments are conducted on four widely used datasets, i.e. N-UCLA, SBU, NTU RGB+D and NTU RGB+D 120 to verify the effectiveness of Ta-CNN. We surpass existing CNN-based methods significantly. Compared with leading GCN-based methods, we achieve comparable performance with much less complexity in terms of the required GFLOPs and parameters.

Kailin Xu, Fanfan Ye, Qiaoyong Zhong, Di Xie• 2021

Related benchmarks

Task	Dataset	Result
Action Recognition	NTU RGB+D 120 (X-set)	Accuracy87.3	779
Action Recognition	NTU RGB+D (Cross-View)	Accuracy94.8	663
Action Recognition	NTU RGB+D 60 (Cross-View)	Accuracy95.1	601
Action Recognition	NTU RGB+D (Cross-subject)	Accuracy90.4	511
Action Recognition	NTU RGB+D 60 (X-sub)	Accuracy90.7	496
Action Recognition	NTU RGB+D X-sub 120	Accuracy86.7	482
Action Recognition	NTU-60 (xsub)	Accuracy90.4	271
Skeleton-based Action Recognition	NTU RGB+D (Cross-View)	Accuracy95.1	213
Skeleton-based Action Recognition	NTU RGB+D 120 (X-set)	Top-1 Accuracy86.8	184
Action Recognition	NTU-60 (xview)	Accuracy94.8	165

Showing 10 of 22 rows

Other info

Follow for update

@wizwand_team Discord