[Paper Review] Music Composition with Deep Learning: A Review
This paper reviews deep learning (DL) approaches to music composition, comparing AI-generated music to human creativity through architectural analysis of VAEs, GANs, LSTMs, and Transformers. It identifies key challenges in structural coherence, creativity evaluation, and style generalization, advocating for improved subjective metrics and interactive, end-to-end models for future research.
Generating a complex work of art such as a musical composition requires exhibiting true creativity that depends on a variety of factors that are related to the hierarchy of musical language. Music generation have been faced with Algorithmic methods and recently, with Deep Learning models that are being used in other fields such as Computer Vision. In this paper we want to put into context the existing relationships between AI-based music composition models and human musical composition and creativity processes. We give an overview of the recent Deep Learning models for music composition and we compare these models to the music composition process from a theoretical point of view. We have tried to answer some of the most relevant open questions for this task by analyzing the ability of current Deep Learning models to generate music with creativity or the similarity between AI and human composition processes, among others.
Motivation & Objective
- To analyze the relationship between AI-based music generation and human compositional creativity.
- To evaluate the effectiveness of deep learning architectures—such as VAEs, GANs, LSTMs, and Transformers—in generating musically coherent and creative compositions.
- To identify open challenges in music generation, including structural coherence, evaluation of creativity, and generalization across musical styles.
- To explore the limitations of current datasets and the need for improved transfer learning and style conditioning in DL models.
- To propose future research directions, including interactive AI-composer collaboration and end-to-end music generation systems.
Proposed method
- Surveyed state-of-the-art deep learning models in music generation, focusing on VAEs, GANs, LSTMs, and Transformers.
- Analyzed architectural components such as latent space learning in VAEs, adversarial training in GANs, and attention mechanisms in Transformers.
- Evaluated model performance using both objective metrics (e.g., log-likelihood, n-gram accuracy) and subjective listening tests.
- Discussed transfer learning strategies using pre-trained NLP models (e.g., Transformers) fine-tuned on symbolic music data.
- Proposed the use of standardized subjective evaluation frameworks—such as MuseGAN’s feature-based rating system—to improve consistency in human evaluation.
- Explored architectural and training design choices that affect long-term musical structure and harmony modeling.
Experimental results
Research questions
- RQ1To what extent can current deep learning models generate music that exhibits structural coherence and harmonic complexity comparable to human compositions?
- RQ2How do the design choices in VAEs, GANs, LSTMs, and Transformers influence the creativity and diversity of generated music?
- RQ3What are the limitations of existing evaluation methods in measuring the creativity and quality of AI-generated music?
- RQ4How can transfer learning and style conditioning be improved to enable generalization beyond the most represented music genres in existing datasets?
- RQ5What role can interactive, human-AI collaborative models play in advancing music generation toward end-to-end, creative composition?
Key findings
- Transformers and NLP-inspired architectures show strong performance in music generation due to their ability to model long-range dependencies in symbolic music sequences.
- VAEs and GANs demonstrate promise in learning disentangled and meaningful latent representations of musical features, though they often struggle with long-term coherence.
- Current objective evaluation metrics are inconsistent across studies, and no unified standard for subjective evaluation exists, limiting comparability.
- Listening tests remain the primary method for assessing perceived quality, but they often fail to capture nuanced aspects of musical creativity and structure.
- Models trained on limited datasets (e.g., JSB Chorales, Lakh MIDI) exhibit poor generalization to underrepresented genres and styles, highlighting a need for diverse training data and better transfer learning.
- Future models should focus on end-to-end generation and interactive design, enabling composers to co-create with AI while maintaining artistic control and originality.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.