Skip to main content
QUICK REVIEW

[Paper Review] Musical Training, but not Mere Exposure to Music, Drives the Emergence of Chroma Equivalence in Artificial Neural Networks

Lukas Grasse, Matthew S. Tata|arXiv (Cornell University)|Feb 20, 2026
Neuroscience and Music Perception0 citations
TL;DR

The study shows chroma equivalence emerges in ANNs only after supervised music transcription fine-tuning, while mere exposure or self-supervised training with music does not; pitch height is more universally represented.

ABSTRACT

Pitch is a fundamental aspect of auditory perception. Pitch perception is commonly described across two perceptual dimensions: pitch height is the sense that tones with varying frequencies seem to be higher or lower, and chroma equivalence is the cyclical similarity of notes octaves, corresponding to a doubling of fundamental frequency. Existing research is divided on whether chroma equivalence is a learned percept that varies according to musical experience and culture, or is an innate percept that develops automatically. Building on a recent framework that proposes to use ANNs to ask 'why' questions about the brain, we evaluated recent auditory ANNs using representational similarity analysis to test the emergence of pitch height and chroma equivalence in their learned representations. Additionally, we fine-tuned two models, Wav2Vec 2.0 and Data2Vec, on a self-supervised learning task using speech and music, and a supervised music transcription task. We found that all models exhibited varying degrees of pitch height representation, but that only models trained on the supervised music transcription task exhibited chroma equivalence. Mere exposure to music through self-supervised learning was not sufficient for chroma equivalence to emerge. This supports the view that chroma equivalence is a higher-order cognitive computation that emerges to support the specific task of music perception, distinct from other auditory perception such as speech listening. This work also highlights the usefulness of ANNs for probing the developmental conditions that give rise to perceptual representations in humans.

Motivation & Objective

  • Investigate whether pitch height and chroma equivalence emerge in ANNs under different training regimes.
  • Determine if self-supervised exposure to music or speech drives chroma equivalence.
  • Assess whether supervised music transcription training is necessary for chroma equivalence to emerge.
  • Compare pretrained, self-supervised, and supervised-finetuned models against chroma and pitch models using RSA.

Proposed method

  • Evaluate transformer-based auditory models (Wav2Vec 2.0, Data2Vec, Whisper, MERT, AST) under SSL, SL, or SSL+SFT.
  • Fine-tune models on speech+music data to test passive music exposure effects on chroma formation.
  • Fine-tune models on polyphonic piano music transcription (MAESTRO) to test active music-task effects on chroma formation.
  • Use Representational Similarity Analysis (RSA) to compare model embeddings to pitch-height and chroma-equivalence models.
  • Stimuli drawn from NSynth notes across octaves 4–6 for RSA with 30 instruments (10 each of flute, guitar, keyboard).
  • Analyze noise ceilings and perform statistical testing with Bonferroni correction.

Experimental results

Research questions

  • RQ1Does pitch height representation emerge in ANNs irrespective of training regime?
  • RQ2Does chroma equivalence emerge with self-supervised exposure to music or speech?
  • RQ3Does supervised fine-tuning on a music transcription task induce chroma equivalence compared to other tasks?
  • RQ4Is mere exposure to music (SSL+ exposure) sufficient for chroma in ANN representations?

Key findings

  • All pretrained/self-supervised ANNs encode pitch height but not chroma equivalence.
  • Incorporating music into training data via self-supervised fine-tuning does not yield chroma equivalence.
  • Supervised fine-tuning on a music transcription task induces chroma equivalence in Wav2Vec 2.0 and Data2Vec.
  • Fine-tuning on speech recognition yields no chroma equivalence gain, despite similar or increased pitch height encoding.
  • CQT-based models show chroma equivalence due to their design, not emergent from general training.
  • Pitch height representations appear broadly automatic, whereas chroma equivalence reflects higher-level music-related computations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.