Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions

About

We demonstrate that carefully adjusting the tokenizer of the Whisper speech recognition model significantly improves the precision of word-level timestamps when applying dynamic time warping to the decoder's cross-attention scores. We fine-tune the model to produce more verbatim speech transcriptions and employ several techniques to increase robustness against multiple speakers and background noise. These adjustments achieve state-of-the-art performance on benchmarks for verbatim speech transcription, word segmentation, and the timed detection of filler events, and can further mitigate transcription hallucinations. The code is available open https://github.com/nyrahealth/CrisperWhisper.

Laurin Wagner, Bernhard Thallinger, Mario Zusag• 2024

Related benchmarks

TaskDatasetResultRank
Automatic Speech RecognitionLibriSpeech (test-other)
WER4
1447
Automatic Speech RecognitionLibriSpeech clean (test)
WER1.82
1410
Automatic Speech RecognitionLibriSpeech Other
WER3.72
140
Automatic Speech RecognitionLibriSpeech Clean
WER1.71
124
Automatic Speech RecognitionAMI
WER8.43
46
Automatic Speech RecognitionVoxPopuli
WER6.03
44
Automatic Speech RecognitionEarnings-22
WER12.9
39
Automatic Speech RecognitionCommon Voice
WER7.76
22
Automatic Speech RecognitionTED-LIUM
WER3.2
22
Automatic Speech RecognitionMLS
WER5.26
7
Showing 10 of 20 rows

Other info

Follow for update