Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Scaling to Multimodal and Multichannel Heart Sound Classification with Synthetic and Augmented Biosignals

About

Cardiovascular diseases (CVDs) are the leading cause of death worldwide, accounting for approximately 17.9 million deaths each year. Early detection is critical, creating a demand for accurate and inexpensive pre-screening methods. Deep learning has recently been applied to classify abnormal heart sounds indicative of CVDs using synchronised phonocardiogram (PCG) and electrocardiogram (ECG) signals, as well as multichannel PCG (mPCG). However, state-of-the-art architectures remain underutilised due to the limited availability of synchronised and multichannel datasets. Augmented datasets and pre-trained models provide a pathway to overcome these limitations, enabling transformer-based architectures to be trained effectively. This work combines traditional signal processing with denoising diffusion models, WaveGrad and DiffWave, to create an augmented dataset to fine-tune a Wav2Vec 2.0-based classifier on multimodal and multichannel heart sound datasets. The approach achieves state-of-the-art performance. On the Computing in Cardiology (CinC) 2016 dataset of single channel PCG, accuracy, unweighted average recall (UAR), sensitivity, specificity and Matthew's correlation coefficient (MCC) reach 92.48%, 93.05%, 93.63%, 92.48%, 94.93% and 0.8283, respectively. Using the synchronised PCG and ECG signals of the training-a dataset from CinC, 93.14%, 92.21%, 94.35%, 90.10%, 95.12% and 0.8380 are achieved for accuracy, UAR, sensitivity, specificity and MCC, respectively. Using a wearable vest dataset consisting of mPCG data, the model achieves 77.13% accuracy, 74.25% UAR, 86.47% sensitivity, 62.04% specificity, and 0.5082 MCC. These results demonstrate the effectiveness of transformer-based models for CVD detection when supported by augmented datasets, highlighting their potential to advance multimodal and multichannel heart sound classification.

Milan Marocchi, Matthew Fynn, Kayapanda Mandana, Yue Rong• 2025

Related benchmarks

TaskDatasetResultRank
Heart Sound ClassificationCinC 2016
UAR92.48
5
Heart Sound ClassificationCinC 2016 (train-a)
Accuracy93.14
4
CAD classificationNoisy PCG dataset Subject Level 1.0
Accuracy77.1
3
CAD classificationNoisy PCG dataset Fragment Level 1.0
Accuracy70.7
3
Abnormal heart sound classificationVest 157 subjects 60s free breathing v1 (96CAD/61NOR)
Accuracy77.13
1
Abnormal heart sound classificationVest 80 subjects 10s breath held v1 (40CAD 40NOR)--
1
Abnormal heart sound classificationVest 157 subjects 10s breath held v1 (96CAD 61NOR)--
1
Showing 7 of 7 rows

Other info

Follow for update