Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

About

Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame rates (e.g., 25 or 12.5 Hz), ignoring the time-varying information density of speech and offering no flexibility to trade off quality for speed at inference time. Recent audio tokenizer research has proposed dynamic frame rate speech coding, which exploits this non-uniformity and enables two new capabilities: very low average frame rates and frame rate controllability. However, this technique has not yet been applied to SLMs. We introduce Flexible Spoken Language Model (FlexiSLM), the first SLM that supports dynamic and controllable frame rates on both speech input and output. Using dynamic frame rate representations, FlexiSLM outperforms fixed-frame-rate 7B models including Qwen2.5-Omni and Kimi-Audio at its high-quality operating points. We further verify that FlexiSLM can be accurately steered down to 4.0 Hz; at 6.25 Hz, it roughly halves inference time relative to 12.5 Hz while retaining strong speech-to-speech quality. Audio samples are available at https://flexislm.github.io .

Jiaqi Li, Chaoren Wang, Xiaohai Tian, Mingjie Chen, Xinyu Liang, Xu Li, Yufan Lin, Junwen Qiu, Jun Zhang, Lu Lu, Haizhou Li, Zhizheng Wu• 2026

Related benchmarks

TaskDatasetResultRank
Automatic Speech RecognitionLibriSpeech (test-other)
WER5.69
1447
Automatic Speech RecognitionLibrispeech (test-clean)
WER1.98
170
Text-to-SpeechLibriSpeech-PC (test)
WER2.14
27
Speech-to-Text reasoningKimi-Audio-Evalkit s2t
Llama Q Score80
18
Speech-to-Speech reasoningKimi-Audio-Evalkit s2s
Llama Q Score75
17
Audio UnderstandingLLaSO-Eval
Accuracy (CremaD Emotion)50
11
Speech GenerationOpenAudioBench (response)
WER4.41
7
Speech GenerationOpenAudioBench (subset)
RTF0.57
5
Showing 8 of 8 rows

Other info

GitHub

Follow for update