Skip to main content
QUICK REVIEW

[论文解读] A variational autoencoder for music generation controlled by tonal tension

Rui Guo, Ivor Simpson|arXiv (Cornell University)|Oct 13, 2020
Music Technology and Sound Studies参考文献 15被引用 11
一句话总结

本文提出一种变分自编码器(VAE)用于可控音乐生成,通过基于螺旋阵列理论整合两种音高张力度量——云直径与张力应变——实现。通过将缩放后的张力方向与张力水平向量引入潜在空间,模型在保持节奏结构的同时,通过调整音高内容以匹配目标张力轮廓,生成符合预期张力特征的音乐变体,从而实现对生成作品中音乐张力的精确控制。

ABSTRACT

Many of the music generation systems based on neural networks are fully autonomous and do not offer control over the generation process. In this research, we present a controllable music generation system in terms of tonal tension. We incorporate two tonal tension measures based on the Spiral Array Tension theory into a variational autoencoder model. This allows us to control the direction of the tonal tension throughout the generated piece, as well as the overall level of tonal tension. Given a seed musical fragment, stemming from either the user input or from directly sampling from the latent space, the model can generate variations of this original seed fragment with altered tonal tension. This altered music still resembles the seed music rhythmically, but the pitch of the notes are changed to match the desired tonal tension as conditioned by the user.

研究动机与目标

  • 开发一种可控音乐生成系统,使用户能够塑造生成音乐中的音高张力特征。
  • 将两种音高张力度量——云直径与张力应变——整合进深度生成模型,实现精细化控制。
  • 在修改音高内容以匹配用户定义的张力轮廓时,保持节奏结构不变。
  • 通过基于表示张力方向与水平的潜在空间向量进行条件化,实现连贯的长时序音乐变体生成。
  • 探究张力控制对生成片段音乐连贯性与感知质量的影响。

提出的方法

  • 模型采用基于张力相关向量条件化的条件变分自编码器(VAE)架构,其潜在空间受张力特征向量调控。
  • 基于螺旋阵列理论,从音乐片段中计算两种音高张力度量——云直径(基于不协和性)与张力应变(基于调性距离)——。
  • 从MIDI文件中提取四小节的单音旋律-低音片段,并将其转至C大调或A小调以保持调性参考一致性。
  • 将张力特征向量(如张力应变方向、云直径水平)进行缩放后加入潜在空间,以调节输出张力。
  • 模型在3,457首流行MIDI文件中的44,900个四小节片段上进行训练,使用Midi-Miner提取音高与张力特征。
  • 生成的音乐保持原始种子片段的节奏结构,但通过调整音高内容以匹配目标张力轮廓。

实验结果

研究问题

  • RQ1VAE模型能否被有效条件化,以生成具有用户定义音高张力轮廓的音乐?
  • RQ2在修改音高内容以匹配张力目标时,模型在多大程度上能保持节奏结构?
  • RQ3不同张力特征向量(如张力应变与云直径)如何影响最终的张力形状与音高分布?
  • RQ4通过依次对潜在表示应用张力向量,模型能否生成连贯的长时序音乐变体?
  • RQ5引入张力感知特征是否能提升生成音乐的感知质量与音高分布?

主要发现

  • 在潜在空间中加入缩放后的张力应变方向向量,可显著提升生成音乐中向上张力趋势的比例,生成样本中向上趋势比例达70%。
  • 云直径方向向量对张力应变向上比例的影响强于张力应变方向向量,表明二者效应存在层级关系。
  • 音高分布发生显著变化:C与G音符频率下降,而A、E与D音符在张力应变向上条件化下显著增强。
  • 通过使用自定义形状的张力向量,模型成功生成具有复杂张力形态(如\diagup\diagdown或\diagdown\diagup)的音乐。
  • 引入张力应变特征可减少低音音高丢失,提升生成音乐的音高准确性。
  • 通过依次对种子片段的潜在空间应用张力向量,系统可实现连贯的多片段音乐生成。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。