[Paper Review] A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions
A survey organizing deep music generation into three levels (score, performance, and audio), detailing representations, datasets, evaluation methods, and future directions.
The utilization of deep learning techniques in generating various contents (such as image, text, etc.) has become a trend. Especially music, the topic of this paper, has attracted widespread attention of countless researchers.The whole process of producing music can be divided into three stages, corresponding to the three levels of music generation: score generation produces scores, performance generation adds performance characteristics to the scores, and audio generation converts scores with performance characteristics into audio by assigning timbre or generates music in audio format directly. Previous surveys have explored the network models employed in the field of automatic music generation. However, the development history, the model evolution, as well as the pros and cons of same music generation task have not been clearly illustrated. This paper attempts to provide an overview of various composition tasks under different music generation levels, covering most of the currently popular music generation tasks using deep learning. In addition, we summarize the datasets suitable for diverse tasks, discuss the music representations, the evaluation methods as well as the challenges under different levels, and finally point out several future directions.
Motivation & Objective
- Classify music generation tasks by representation level (score, performance, audio) to enable targeted literature retrieval.
- Summarize music representations, datasets, and evaluation methodologies used across deep learning approaches.
- Analyze the strengths and limitations of deep learning models for different music generation tasks.
- Highlight challenges and propose future research directions in deep music generation.
Proposed method
- Review and categorize existing work according to music generation levels and tasks.
- Compare symbolic (MIDI-like) and audio representations, including REMI and tuple-based encodings.
- Survey deep learning architectures applied to score, performance, and audio generation (RNNs, LSTMs, VAEs, GANs, Transformers).
- Discuss datasets, evaluation methods (objective and subjective), and cross-modal/fusion possibilities.
- Summarize historical approaches and dataset availability to contextualize modern deep learning methods.
Experimental results
Research questions
- RQ1How are music generation tasks organized across score, performance, and audio representations in deep learning work?
- RQ2What representations, datasets, and evaluation methods dominate each generation level?
- RQ3What are the key strengths and limitations of current models (RNNs, VAEs, GANs, Transformers) for different generation tasks?
- RQ4What future directions and challenges are most impactful for advancing deep music generation?
Key findings
- Deep learning architectures have become mainstream in music generation, spanning score, performance, and audio tasks.
- REMI and other representation schemes improve rhythm modeling over traditional MIDI-like representations.
- Transformers and hierarchical VAEs are influential for long-term structure capture and cross-modal generation.
- WaveNet and GAN-based approaches significantly advance audio synthesis and singing voice synthesis.
- A lack of aligned score-performance datasets is a bottleneck for expressive performance modeling.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.