Skip to main content
QUICK REVIEW

[论文解读] Using perceptually defined music features in music information retrieval

Anders Friberg, Erwin Schoonderwaldt|arXiv (Cornell University)|Mar 31, 2014
Music and Audio Processing参考文献 39被引用 6
一句话总结

本文提出了一种基于人类听觉感知而非传统音乐理论的新型音乐信息检索框架,通过定义感知音乐特征来实现。利用MIDI和音频数据的听觉实验与计算建模,结果表明,少量感知特征可预测高达90%的情绪表达方差,优于传统的音频特征集。

ABSTRACT

In this study, the notion of perceptual features is introduced for describing general music properties based on human perception. This is an attempt at rethinking the concept of features, in order to understand the underlying human perception mechanisms. Instead of using concepts from music theory such as tones, pitches, and chords, a set of nine features describing overall properties of the music was selected. They were chosen from qualitative measures used in psychology studies and motivated from an ecological approach. The selected perceptual features were rated in two listening experiments using two different data sets. They were modeled both from symbolic (MIDI) and audio data using different sets of computational features. Ratings of emotional expression were predicted using the perceptual features. The results indicate that (1) at least some of the perceptual features are reliable estimates; (2) emotion ratings could be predicted by a small combination of perceptual features with an explained variance up to 90%; (3) the perceptual features could only to a limited extent be modeled using existing audio features. The results also clearly indicated that a small number of dedicated features were superior to a 'brute force' model using a large number of general audio features.

研究动机与目标

  • 通过将音乐特征建立在人类感知基础上,而非音乐理论,重新定义音乐特征。
  • 识别一组最小化的感知特征,以可靠地描述音乐的一般属性。
  • 评估这些特征在预测音乐情绪表达方面的预测能力。
  • 比较感知特征与标准音频特征在建模情绪评分方面的有效性。

提出的方法

  • 从心理学研究中选取九种感知特征,基于生态感知框架。
  • 在两个独立的音乐数据集上开展两次听觉实验,对感知特征进行评分。
  • 使用计算特征集,从符号化(MIDI)和音频数据中建模感知特征。
  • 使用回归模型,基于感知特征预测情绪表达评分。
  • 将感知特征模型的性能与使用大量通用音频特征的“暴力”模型进行比较。
  • 使用决定系数(R²)作为主要指标,评估模型性能。

实验结果

研究问题

  • RQ1能否通过人类感知推导出的感知特征可靠地描述音乐的一般属性?
  • RQ2仅使用少量感知特征,能在多大程度上预测音乐中的情绪表达?
  • RQ3感知特征在预测情绪评分方面与标准音频特征相比表现如何?
  • RQ4与大量通用音频特征相比,有针对性的感知特征集在情绪预测中是否更有效?

主要发现

  • 部分感知特征在听觉实验中表现出可靠的评分者间一致性。
  • 仅使用少量感知特征的组合,即可实现高达90%的方差解释率来预测情绪表达评分。
  • 现有音频特征只能部分建模感知特征,表明当前音频特征表示存在缺陷。
  • 少量专门设计的感知特征显著优于使用大量通用音频特征的“暴力”方法。
  • 结果支持基于感知的特征作为传统音频特征集在音乐情绪预测中的更有效替代方案。
  • 本研究表明,感知特征更能捕捉到驱动情绪反应的心理学维度。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。