MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
About
We present MIDI-LLM, an LLM for generating multitrack MIDI music from free-form text prompts. Our approach expands a text LLM's vocabulary to include MIDI tokens, and uses a two-stage training recipe to endow text-to-MIDI abilities. By preserving the original LLM's parameter structure, we can directly leverage the vLLM library for accelerated inference. Experiments show that MIDI-LLM achieves higher quality, better text control, and faster inference compared to the recent Text2midi model. Live demo at https://midi-llm-demo.vercel.app.
Shih-Lun Wu, Yoon Kim, Cheng-Zhi Anna Huang• 2025
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Text-to-score generation | Custom Text-to-Score prompt set 1.0 (test) | Valid File Generation Rate100 | 5 | |
| Text-to-score generation | 238 evaluation prompts | Prompt Adherence1.67 | 3 |
Showing 2 of 2 rows