Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Fisher

Benchmarks

Task NameDataset NameSOTA ResultTrend
Conditional Dialogue GenerationFisher (test)
GPT-4o Score7.3
64
Speech-to-speech translationFisher Spanish-English (test)
BLEU (Speech Input)90.5
55
Speech-to-speech translationFisher Spanish-English (dev)
BLEU (Speech)88.5
48
Speech-to-speech translationFisher Spanish-English (dev2)
ASR BLEU89.4
36
Unconditional Dialogue GenerationFisher (test)
GPT-4o Score9.33
32
Speech TranslationFisher Monolingual (test)
BLEU35.87
11
Speech TranslationFisher Code-Switching (test)
BLEU37.51
11
Speaker-Attributed Automatic Speech RecognitionFisher (test)
WDER0.9
11
Speech-to-Speech TranslationFisher Es→En (test)
ASR chrF70.2
10
Speech-to-Speech TranslationFisher Es→En (dev)
ASR chrF69.5
10
Conditional Turn-taking EvaluationFisher (test)
Occurrence Proportion58
7
Unconditional Turn-taking EvaluationFisher (test)
Occurrence Rate60
7
BackchannelingFisher
Init Rate97.8
5
Window-level Turn-takingFisher
Onset MAE0.69
5
Dialogue GenerationFisher
M-MOS4.25
4
ES-to-EN ASTFisher (test)
BLEU64.7
4
Speaker-Attributed Automatic Speech RecognitionFisher Global Meeting-level
DER15.21
4
Speaker-Attributed Automatic Speech RecognitionFisher (local setting)
DER8.18
4
Cross-lexical backchannel similarityFisher
Proportion of Correct Selections66.3
3
Prosodic backchannel similarityFisher
Proportion Correct Selections69.7
3
Fine-grained Score AccuracyFisher
Exact Accuracy64.76
1
Binary classification (Human vs Machine speech)Fisher (Human-Human) OOD (test)
Accuracy98.44
1
Showing 22 of 22 rows