FBK's Long-form SpeechLLMs for IWSLT 2026 Instruction Following
About
This paper describes our submission to the IWSLT 2026 Instruction Following shared task. SpeechLLMs are developed for both short-form and long-form speech instruction following under constrained settings. For the short track, strong performance is achieved on MCIF, with a SIFS score of 2.0708. For the long track, three speech segmentation methods are explored, and the HIFS score is introduced to account for unstable long-form generation. Experimental results show that fixed 30-second segmentation provides the most robust long-form performance, achieving the highest HIFS score of 2.0663. Further analysis shows that hallucination mainly manifests as repetitive insertions in generated outputs, substantially affecting ASR and SSUM, while short-form capabilities are largely retained after long-form extension.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Speech Question Answering | MCIF short-form | BERTScore (SQA)45.13 | 10 | |
| Speech Translation | MCIF short-form | ST COMET78.69 | 8 | |
| Automatic Speech Recognition | MCIF short-form | ASR Accuracy88.77 | 4 | |
| ACHAP Track | IWSLT long-form EN-EN 2026 | WER20 | 3 | |
| ACHAP Track | IWSLT long-form EN-DE 2026 | COMET0.69 | 3 | |
| ACHAP Track | IWSLT long-form EN-IT 2026 | COMET0.735 | 3 | |
| ACHAP Track | IWSLT long-form EN-ZH 2026 | COMET0.698 | 3 | |
| Automatic Speech Recognition | IWSLT long-form EN-EN 2026 | WER12.6 | 3 | |
| Quality Estimation | IWSLT long-form EN-DE 2026 | Accuracy50.1 | 3 | |
| Quality Estimation | IWSLT long-form EN-ZH 2026 | Accuracy65.8 | 3 |