Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Aligning MusicLLM with Emotion using Instruction Tuning and Feedback-Driven Alignment

About

This paper investigates whether music large language models (MusicLLMs) can be aligned for emotion regression. While MusicLLMs have shown strong performance in music information retrieval tasks, their ability to predict arousal and valence scores remains limited, since emotion regression has not been an explicit training objective. To examine whether MusicLLMs can be aligned with emotion, we train MusicLLMs on emotion regression and compare two strategies: instruction tuning and feedback-driven alignment. Our experiments show that task-aware instruction tuning enables MusicLLMs to predict emotion levels to some extent, although the accuracy remains limited. Applying feedback-driven alignment with a verifiable numerical reward substantially improves performance on both arousal and valence over instruction tuning alone. We further show that our approach improves emotion regression performance while maintaining MusicQA capability.

Takuya Hasumi, Welly Naptali• 2026

Related benchmarks

TaskDatasetResultRank
Music Emotion RegressionDEAM
R^2 (Arousal)0.48
5
Music Emotion RegressionMerge
R^2 (Arousal)0.5
5
Music Question AnsweringMusicQA
BLEU@40.15
5
Music Emotion RegressionMERGE (test)
R^2 (Arousal)0.55
4
Music Emotion RegressionDEAM (test)
R^2 (Arousal)0.56
4
Showing 5 of 5 rows

Other info

Follow for update