[论文解读] A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions
一份将深度音乐生成分为三个层级(score, performance, and audio)的综述,详细介绍表示、数据集、评估方法和未来方向。
The utilization of deep learning techniques in generating various contents (such as image, text, etc.) has become a trend. Especially music, the topic of this paper, has attracted widespread attention of countless researchers.The whole process of producing music can be divided into three stages, corresponding to the three levels of music generation: score generation produces scores, performance generation adds performance characteristics to the scores, and audio generation converts scores with performance characteristics into audio by assigning timbre or generates music in audio format directly. Previous surveys have explored the network models employed in the field of automatic music generation. However, the development history, the model evolution, as well as the pros and cons of same music generation task have not been clearly illustrated. This paper attempts to provide an overview of various composition tasks under different music generation levels, covering most of the currently popular music generation tasks using deep learning. In addition, we summarize the datasets suitable for diverse tasks, discuss the music representations, the evaluation methods as well as the challenges under different levels, and finally point out several future directions.
研究动机与目标
- 将音乐生成任务按表示层级(score, performance, audio)分类,以实现有针对性的文献检索。
- 总结跨深度学习方法使用的音乐表示、数据集和评估方法。
- 分析深度学习模型在不同音乐生成任务中的优缺点。
- 突出挑战并提出深度音乐生成的未来研究方向。
提出的方法
- 评审并按照音乐生成的层级和任务对现有工作进行分类。
- 比较符号(MIDI 类)表示和音频表示,包括 REMI 和基于元组的编码。
- 调查应用于 score、performance、和 audio 生成的深度学习架构(RNN、LSTM、VAE、GAN、Transformer)。
- 讨论数据集、评估方法(客观和主观),以及跨模态/融合的可能性。
- 总结历史方法和数据集可用性,以为现代深度学习方法提供背景。
实验结果
研究问题
- RQ1在深度学习工作中,音乐生成任务如何在 score、performance、和 audio 表征之间组织?
- RQ2每个生成层级主导的表示、数据集和评估方法是什么?
- RQ3当前模型(RNNs、VAEs、GANs、Transformers)在不同生成任务中的关键优缺点是什么?
- RQ4哪些未来方向和挑战对推动深度音乐生成最具影响力?
主要发现
- 深度学习架构已经在音乐生成中成为主流,涵盖 score、performance、和 audio 任务。
- REMI 及其他表示方案在节奏建模方面优于传统的 MIDI-like 表示。
- Transformers 与分层 VAE 对长时结构捕捉和跨模态生成具有影响力。
- WaveNet 和基于 GAN 的方法显著推动音频合成和人声合成。
- 缺乏对齐的乐谱-演奏数据集是表达性演奏建模的瓶颈。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。