[论文解读] From Bach to the Beatles: The simulation of human tonal expectation using ecologically-trained predictive models
该论文表明,经过真实世界音乐录音(如巴赫的《平均律键盘曲集》和披头士乐队作品)的原始音频训练的循环神经网络(RNNs)会发展出与人类探测音评级高度一致的调性预期,表明其隐式学习到了调性结构。关键发现是,尽管调性结构对于准确的音乐预期是必要的,但并不充分,因为节奏和声部进行等时间性线索也发挥着关键作用。
Tonal structure is in part conveyed by statistical regularities between musical events, and research has shown that computational models reflect tonal structure in music by capturing these regularities in schematic constructs like pitch histograms. Of the few studies that model the acquisition of perceptual learning from musical data, most have employed self-organizing models that learn a topology of static descriptions of musical contexts. Also, the stimuli used to train these models are often symbolic rather than acoustically faithful representations of musical material. In this work we investigate whether sequential predictive models of musical memory (specifically, recurrent neural networks), trained on audio from commercial CD recordings, induce tonal knowledge in a similar manner to listeners (as shown in behavioral studies in music perception). Our experiments indicate that various types of recurrent neural networks produce musical expectations that clearly convey tonal structure. Furthermore, the results imply that although implicit knowledge of tonal structure is a necessary condition for accurate musical expectation, the most accurate predictive models also use other cues beyond the tonal structure of the musical context.
研究动机与目标
- 探究生态化训练的预测模型是否能够模拟行为研究中观察到的人类调性预期。
- 确定在原始音频而非符号化表示上训练的循环神经网络(RNNs)是否能通过暴露隐式学习调性结构。
- 评估调性结构是否为计算模型中准确音乐预期的必要条件。
- 探讨除调性之外的非调性线索(如节奏、声部进行和终止式结构)在塑造音乐预测中的作用。
提出的方法
- 在巴赫《平均律键盘曲集》和披头士乐队商业CD录音的原始音频频谱图(CQT)上训练多种类型的循环神经网络(RNNs)。
- 采用预测编码原理:模型通过最小化基于当前上下文对未来音乐事件的预测误差进行训练。
- 通过与人类探测音评级(Krumhansl & Kessler, 1982)的相关性来评估模型的预期,使用皮尔逊相关系数作为度量指标。
- 通过比较模型在原始音频与时间顺序打乱音频数据上的表现,分离出序列结构的作用。
- 使用平均交叉熵损失衡量预测准确性,使用与人类探测音轮廓的平均相关性衡量调性结构。
- 分析不同音乐类型和结构复杂度(如单音钢琴曲与多乐器流行音乐)下模型的行为。
实验结果
研究问题
- RQ1在真实世界音频录音上训练的RNN是否能以类似于人类听者的方式学习调性结构?
- RQ2调性结构是否为预测模型中准确音乐预期的必要条件?
- RQ3非调性线索(如节奏、声部进行和终止式结构)在调性结构之外对音乐预测的贡献程度如何?
- RQ4在时间顺序打乱的音频上进行训练如何影响模型性能和调性预期,这揭示了序列顺序的何种作用?
主要发现
- 在披头士乐队和巴赫《平均律键盘曲集》的原始音频上训练的RNN产生的音乐预期与人类探测音评级高度相关,表明成功模拟了人类的调性层级结构。
- 在打乱音频数据上训练的模型表现出显著降低的预测准确性和更弱的调性结构表征,尤其是在巴赫《平均律键盘曲集》中表现更明显,证明了时间顺序在学习中的重要性。
- 调性结构是准确音乐预期的必要条件,因为预测误差低的模型始终与人类评分呈现高度相关性。
- 然而,仅靠调性结构本身不足以实现准确预测,因为即使调性相关性高的模型仍可能具有较高的预测误差,表明节奏和声部进行等额外线索存在显著影响。
- 数据打乱对巴赫《平均律键盘曲集》的影响比对披头士乐队录音更为显著,可能是因为巴赫作品具有更高的和声与旋律复杂性以及频繁的转调。
- 即使在音乐结构更简单的多乐器流行录音(如披头士乐队)中,RNN仍能有效学习调性结构,表明生态上合理的训练数据可支持稳健的感知学习。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。