Skip to main content
QUICK REVIEW

[论文解读] Music Style Transfer Issues: A Position Paper.

Shuqi Dai, Zheng Zhang|arXiv (Cornell University)|Mar 19, 2018
Music and Audio Processing参考文献 31被引用 18
一句话总结

本文认为,由于音乐具有多层次、多模态的表征特性,音乐风格迁移缺乏科学基础,这与图像数据有本质不同。本文提出应将风格迁移重新定义为与成熟的计算机音乐子领域对齐,并利用无监督解耦技术,以克服简单端到端神经网络迁移方法的局限性。

ABSTRACT

Led by the success of neural transfer on visual arts, there has been a rising trend very recently in the effort of transfer. However, music style is not yet a well-defined concept from a scientific point of view. The difficulty lies in the intrinsic multi-level and multi-modal character of representation (which is very different from image representation). As a result, depending on their interpretation of music style, current studies under the category of music are actually solving completely different problems that belong to a variety of sub-fields of Computer Music. Also, a vanilla end-to-end approach, which aims at dealing with all levels of representation at once by directly adopting the method of image transfer, leads to poor results. Thus, we see a vital necessity to re-define transfer more precisely and scientifically based on the uniqueness of representation, as well as to connect different aspects of transfer with existing well-established sub-fields of computer studies. Otherwise, an accumulated upcoming literature (all named after transfer) will lead to a great confusion of the underlying problems as well as negligence of the treasures in computer before the age of deep learning. In addition, we discuss the current limitations of modeling and its future directions by drawing spirit from some deep generative models, especially the ones using unsupervised learning and disentanglement techniques.

研究动机与目标

  • 解决当前风格迁移研究中‘音乐风格’缺乏科学定义的问题。
  • 指出当前音乐风格迁移方法因风格定义模糊,实际上在不同计算机音乐子领域中解决的是无关问题。
  • 批判将直接基于图像的神经网络迁移方法应用于音乐时的失败,原因在于音乐的复杂表征方式。
  • 主张为音乐的独特结构与感知层次,提出一种精确且科学的风格迁移重新定义。
  • 将风格迁移研究与现有、成熟的计算机音乐学科领域连接,以避免混淆并重新发现深度学习出现前的洞见。

提出的方法

  • 基于音乐表征的内在多层次与多模态特性,重新定义音乐风格迁移,而非直接套用图像风格迁移方法。
  • 建议将音乐风格迁移与计算机音乐中已确立的子领域对齐,如音频信号处理、音乐信息检索和符号音乐生成。
  • 引入使用无监督深度生成模型结合解耦技术,以分别建模不同的音乐属性(例如音色、和声、节奏)。
  • 倡导采用分层建模方法,在不同抽象层次上分离风格成分(例如,音符级、和弦级、乐句级)。
  • 通过将风格迁移建立在领域特定知识基础上,而非依赖通用神经网络架构,来提升方法论的严谨性。
  • 从解耦表征学习中汲取灵感,以提升音乐风格迁移中的可解释性与控制性。

实验结果

研究问题

  • RQ1鉴于音乐风格具有多层次与多模态的特性,如何以科学严谨的方式定义‘音乐风格’?
  • RQ2为何简单的端到端神经网络迁移方法在应用于音乐时会失败,而这类方法在图像迁移中却有效?
  • RQ3如何将音乐风格迁移与现有、成熟的计算机音乐子领域有意义地连接起来?
  • RQ4无监督解耦技术在提升音乐风格迁移的可解释性与质量方面能发挥什么作用?
  • RQ5在深度学习兴起的背景下,如何通过重新定义风格迁移来避免混淆与先验知识的丢失?

主要发现

  • 当前的音乐风格迁移方法并非统一,而是因风格定义模糊,在不同计算机音乐子领域中解决的是不同问题。
  • 将基于图像的神经网络迁移方法直接应用于音乐会导致性能低下,原因在于音乐的表征方式与图像有本质不同。
  • 音乐风格缺乏精确的定义,导致文献中出现不一致且缺乏科学依据的研究主张。
  • 将风格迁移与成熟的计算机音乐学科重新连接,可防止重要深度学习前洞见的丢失,并提升方法论的清晰度。
  • 无监督解耦技术为分别建模不同的音乐属性提供了有前景的路径,从而增强控制性与可解释性。
  • 必须重新定义一种科学基础扎实的风格迁移方法,以避免混淆,并确保该领域实现有意义的进展。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。