Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Voice Activity Projection: Self-supervised Learning of Turn-taking Events

About

The modeling of turn-taking in dialog can be viewed as the modeling of the dynamics of voice activity of the interlocutors. We extend prior work and define the predictive task of Voice Activity Projection, a general, self-supervised objective, as a way to train turn-taking models without the need of labeled data. We highlight a theoretical weakness with prior approaches, arguing for the need of modeling the dependency of voice activity events in the projection window. We propose four zero-shot tasks, related to the prediction of upcoming turn-shifts and backchannels, and show that the proposed model outperforms prior work.

Erik Ekstedt, Gabriel Skantze• 2022

Related benchmarks

TaskDatasetResultRank
Turn-taking Naturalness EvaluationTurn-taking Perturbation Benchmark 1.0 (test)
Delta m_theta0.36
10
Endpoint AnticipationSpokenWoz (test)
MRA (ms)80
9
Agent action predictionSwitchboard 138-session (test)
wF138.9
9
Active Shift-Hold PredictionAVCC 2 Speakers (test)
F1 Score62.2
4
Silent Shift-Hold PredictionAVCC 2 Speakers (test)
F1 Score67.2
4
Silent Shift-Hold PredictionAVCC 3 Speakers (test)
F1 Score65.5
4
3-class audio segmentation (Gap/SU/Pause)OpenETD synthetic 1.0 (test)
F1 Score90.6
4
Active Shift-Hold PredictionAVCC 3 Speakers (test)
F1 Score63.4
4
3-class audio segmentation (Gap/SU/Pause)OpenETD real 1.0 (test)
F1 Score33.2
4
Binary ClassificationOpenETD Synthetic
Precision91.5
3
Showing 10 of 13 rows

Other info

Follow for update