Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition

About

Underwater acoustic target recognition (UATR) plays a vital role in marine applications but remains challenging due to limited labeled data and the complexity of ocean environments. This paper explores a central question: can speech large models (SLMs), trained on massive human speech corpora, be effectively transferred to underwater acoustics? To investigate this, we propose UATR-SLM, a simple framework that reuses the speech feature pipeline, adapts the SLM as an acoustic encoder, and adds a lightweight classifier.Experiments on the DeepShip and ShipsEar benchmarks show that UATR-SLM achieves over 99% in-domain accuracy, maintains strong robustness across variable signal lengths, and reaches up to 96.67% accuracy in cross-domain evaluation. These results highlight the strong transferability of SLMs to UATR, establishing a promising paradigm for leveraging speech foundation models in underwater acoustics.

Mengcheng Huang, Xue Zhou, Chen Xu, Dapeng Man• 2026

Related benchmarks

TaskDatasetResultRank
Acoustic Target RecognitionShipsEar (test)
Accuracy99.09
15
Acoustic Target RecognitionDeepShip (test)
Accuracy99.32
7
Showing 2 of 2 rows

Other info

Follow for update