Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users

About

Voice control offers an intuitive alternative to manual drone piloting, yet most existing systems rely on rigid command vocabularies that fail to handle the spontaneous, disfluent speech of naive users. This paper addresses this gap by proposing an End-to-End Spoken Language Understanding architecture for real-time human-drone interaction in French. Our model combines a frozen Self-Supervised Learning acoustic encoder with a lightweight LSTM-based classification head, augmented by a cross-modal knowledge distillation objective that aligns acoustic representations with semantic embeddings from a text teacher, without requiring transcription at inference time. We evaluate our approach on VoiceStick, a novel French corpus of spontaneous speech collected during real teleoperation sessions with 29 nonexpert dyads. On simple voice commands, our best configuration achieves 93% accuracy at 7 ms inference latency, outperforming cascade baselines (79%, 202 ms) with a 29x speedup. On the full spontaneous speech test set, our architecture reaches 82% accuracy, with crossmodal distillation consistently improving robustness across all configurations. These results demonstrate that End-to-End architectures are not only feasible but preferable for spontaneous voice-guided UAV teleoperation, combining semantic robustness, low latency, and calibrated confidence.

Allan Henry, Solange Rossato, Christian Graff, Sylvain Huet, Jose-Ernesto Gomez-Balderas• 2026

Related benchmarks

TaskDatasetResultRank
Intent RecognitionVoiceStick Explicit Subset
Accuracy93
11
Intent RecognitionVoiceStick (Full Set)
Accuracy86
11
Showing 2 of 2 rows

Other info

Follow for update