[Paper Review] A Framework for Automated Pop-song Melody Generation with Piano Accompaniment Arrangement
This paper presents a unified framework for automated pop-song melody and piano accompaniment generation, using chord progressions as input. It employs a harmony alternation model, a seasonal ARMA-based melody generation model, and a melody integration model to produce musically coherent results, achieving subjective ratings significantly higher than bi-LSTM baselines and comparable performance to commercial software like Band in a Box.
We contribute a pop-song automation framework for lead melody generation and accompaniment arrangement. The framework reflects the major procedures of human music composition, generating both lead melody and piano accompaniment by a unified strategy. Specifically, we take chord progression as an input and propose three models to generate a structured melody with piano accompaniment textures. First, the harmony alternation model transforms a raw input chord progression to an altered one to better fit the specified music style. Second, the melody generation model generates the lead melody and other voices (melody lines) of the accompaniment using seasonal ARMA (Autoregressive Moving Average) processes. Third, the melody integration model integrates melody lines (voices) together as the final piano accompaniment. We evaluate the proposed framework using subjective listening tests. Experimental results show that the generated melodies are rated significantly higher than the ones generated by bi-directional LSTM, and our accompaniment arrangement result is comparable with a state-of-the-art commercial software, Band in a Box.
Motivation & Objective
- To develop an automated framework that mirrors human music composition processes for pop songs.
- To generate both lead melody and piano accompaniment in a unified, structured manner.
- To improve melody quality and accompaniment coherence using statistical modeling and harmony refinement.
- To evaluate the framework's performance through subjective listening tests against existing methods.
Proposed method
- The harmony alternation model modifies raw chord progressions to better fit a target music style using style-aware transformations.
- The melody generation model uses seasonal ARMA processes to generate lead melody and additional voices for the accompaniment.
- The melody integration model combines multiple generated voices into a final, playable piano accompaniment texture.
- The framework processes chord progressions as input and outputs a complete, harmonically coherent pop song with melody and accompaniment.
- All components are trained and evaluated within a single, end-to-end pipeline to ensure musical consistency.
- The system is evaluated via subjective listening tests to assess perceptual quality and musicality.
Experimental results
Research questions
- RQ1Can a unified framework generate musically coherent pop-song melodies and piano accompaniments from chord progressions?
- RQ2How does the use of seasonal ARMA processes improve melody generation compared to recurrent models like bi-LSTM?
- RQ3To what extent does the harmony alternation model enhance stylistic fitting of chord progressions?
- RQ4How does the integrated output compare to commercial software in terms of perceived quality?
- RQ5What is the relative impact of each component (harmony, melody, integration) on the final musical output?
Key findings
- The generated melodies received significantly higher ratings in subjective listening tests compared to those produced by a bi-directional LSTM baseline.
- The framework's accompaniment arrangements were rated as comparable to those of the state-of-the-art commercial software Band in a Box.
- The harmony alternation model successfully improved chord progression suitability for the target music style, enhancing overall musical coherence.
- The seasonal ARMA-based melody generation model produced more rhythmically and melodically structured outputs than recurrent baselines.
- The melody integration model effectively combined multiple voices into a balanced, playable piano texture without dissonance or voice collision.
- The full framework demonstrated strong musicality and coherence, validating the effectiveness of its modular, unified design.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.