Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SF-Flow: Sound field magnitude estimation via flow matching guided by sparse measurements

About

Reconstructing a 3D sound field from sparse microphone measurements is a fundamental yet ill-posed problem, which we address through Acoustic Transfer Function (ATF) magnitude estimation. ATF magnitude encapsulates key perceptual and acoustic properties of a physical space with applications in room characterization and correction. Although recent generative paradigms such as Flow Matching (FM) have achieved state-of-the-art performance in speech and music generation, their potential in spatial audio remains underexplored. We propose a novel framework for 3D ATF magnitude reconstruction as a guided generation task, with a 3D U-Net conditioned by a permutation-invariant set encoder. This architecture enables reconstruction from an arbitrary number of sparse inputs while leveraging the stable and efficient training properties of FM. Experimental results demonstrate that SF-Flow achieves accurate reconstruction up to \SI{1}{kHz}, trains substantially faster than the autoencoder baseline, and improves significantly with dataset size.

Ege Erdem, Shoichi Koyama, Tomohiko Nakamura, Orchisama Das, Zoran Cvetkovi\'c• 2026

Related benchmarks

TaskDatasetResultRank
3D ATF reconstructionR1 (test)
LSD (0-20 / 312 Hz)1.76
3
3D ATF reconstructionR3 (test)
LSD (0-20 / 312 Hz)0.55
2
3D ATF reconstructionR2 (test)
LSD (0-20 / 312 Hz)0.78
1
Showing 3 of 3 rows

Other info

Follow for update