Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Speaker head orientation estimation with a single microphone array using phase spectrogram features

About

Estimating a speaker's head orientation from audio can provide valuable information in smart environments, meetings, and driver monitoring. We propose a novel approach that leverages the phase component of the short-time Fourier transform from a single microphone array as input to a deep neural network combining convolutional, recurrent, and self-attention layers. Unlike prior methods that use physics-informed handcrafted features or raw waveform inputs, our approach enables robust learning from simulated and real data. Trained on a large-scale dataset generated with voice directivity patterns and fine-tuned on real recordings, our model achieves state-of-the-art accuracy, outperforming baselines under both clean and noisy conditions. Personalization experiments further demonstrate significant gains, reaching a mean angular error of 11.3 degrees when adapting to individual users and environments.

Balint Turi, Archontis Politis, Parthasaarathy Sudarsanam, Tuomas Virtanen• 2026

Related benchmarks

TaskDatasetResultRank
Orientation EstimationSimulated dataset Anechoic 1.0 (test)
MAE (degrees)28.4
3
Orientation EstimationSimulated dataset Clean 1.0 (test)
Mean Absolute Error (MAE)19.9
3
Orientation EstimationSimulated dataset Moderate noise 10–20 dB SNR 1.0 (test)
MAE25.6
3
Orientation EstimationSimulated dataset High noise 0–10 dB SNR 1.0 (test)
MAE29.5
3
Speaker head orientation estimationReal dataset (Session cross-evaluation)
Accuracy73.2
3
Showing 5 of 5 rows

Other info

Follow for update