Publications
CONFERENCE (INTERNATIONAL) Aligning MusicLLM with Emotion using Instruction Tuning and Feedback-Driven Alignment
Takuya Hasumi, Welly Naptali
The 27th Annual Conference of the International Speech Communication Association (INTERSPEECH 2026)
September 29, 2026
This paper investigates whether music large language models (MusicLLMs) can be aligned for emotion regression. While MusicLLMs have shown strong performance in music information retrieval tasks, their ability to predict arousal and valence scores remains limited, since emotion regression has not been an explicit training objective. To examine whether MusicLLMs can be aligned with emotion, we train MusicLLMs on emotion regression and compare two strategies: instruction tuning and feedback-driven alignment. Our experiments show that task-aware instruction tuning enables MusicLLMs to predict emotion levels to some extent, although the accuracy remains limited. Applying feedback-driven alignment with a verifiable numerical reward substantially improves performance on both arousal and valence over instruction tuning alone. We further show that our approach improves emotion regression performance while maintaining MusicQA capability.
Paper :
Aligning MusicLLM with Emotion using Instruction Tuning and Feedback-Driven Alignment
(external link)
PDF : Aligning MusicLLM with Emotion using Instruction Tuning and Feedback-Driven Alignment