Skip to main content
QUICK REVIEW

[Paper Review] A Framework for Multi-f0 Modeling in SATB Choir Recordings

Helena Cuesta, Emília Gómez|arXiv (Cornell University)|Apr 10, 2019
Music and Audio Processing15 references4 citations
TL;DR

This paper proposes a two-stage framework for modeling fundamental frequency (f0) in SATB choir recordings, combining deep learning-based multi-f0 estimation with spectral whitening and peak detection to model f0 distribution per choir section. The key contribution is capturing pitch dispersion (in cents) as a descriptor of unison singing quality, showing consistent trends with prior ground-truth studies despite limited absolute accuracy.

ABSTRACT

Fundamental frequency (f0) modeling is an important but relatively unexplored aspect of choir singing. Performance evaluation as well as auditory analysis of singing, whether individually or in a choir, often depend on extracting f0 contours for the singing voice. However, due to the large number of singers, singing at a similar frequency range, extracting the exact individual pitch contours from choir recordings is a challenging task. In this paper, we address this task and develop a methodology for modeling pitch contours of SATB choir recordings. A typical SATB choir consists of four parts, each covering a distinct range of pitches and often with multiple singers each. We first evaluate some state-of-the-art multi-f0 estimation systems for the particular case of choirs with a single singer per part, and observe that the pitch of individual singers can be estimated to a relatively high degree of accuracy. We observe, however, that the scenario of multiple singers for each choir part (i.e. unison singing) is far more challenging. In this work we propose a methodology based on combining a multi-f0 estimation methodology based on deep learning followed by a set of traditional DSP techniques to model f0 and its dispersion instead of a single f0 trajectory for each choir part. We present and discuss our observations and test our framework with different singer configurations.

Motivation & Objective

  • To address the challenge of modeling individual pitch contours in SATB choir recordings where multiple singers sing in unison within similar pitch ranges.
  • To improve f0 estimation accuracy for choral recordings by combining deep learning with traditional DSP techniques.
  • To model pitch dispersion (variability) across choir sections as a key descriptor of intonation quality in unison singing.
  • To evaluate the performance of state-of-the-art multi-f0 estimation systems on choral data and identify limitations in multi-singer scenarios.
  • To provide a robust, scalable method for analyzing choir intonation without requiring individual vocal track separation.

Proposed method

  • First, a deep learning-based multi-f0 estimation system (DeepSalience) is applied to extract initial f0 candidates per choir section.
  • Second, the input audio signal is spectrally whitened to enhance frequency resolution and reduce spectral tilt.
  • Peaks in the whitened spectrum are detected to locate the dominant f0 energy distribution for each choir section.
  • The mean f0 and bandwidth (as a measure of dispersion) are computed from the detected peaks, representing the f0 distribution of a unison section.
  • The framework uses a two-stage approach: initial f0 estimation followed by high-resolution spectral analysis to model dispersion.
  • The method is evaluated across different choir configurations (1 vs. 4 singers per section) to assess robustness and consistency.

Experimental results

Research questions

  • RQ1How do state-of-the-art multi-f0 estimation systems perform on SATB choir recordings with multiple singers per part?
  • RQ2Can a combination of deep learning and traditional DSP techniques improve f0 resolution and dispersion modeling in unison choir singing?
  • RQ3Does the proposed framework reliably capture pitch dispersion trends across choir sections (soprano, alto, tenor, bass) compared to ground-truth data?
  • RQ4How does the number of singers per choir section affect the estimated f0 dispersion?
  • RQ5Can f0 dispersion serve as a meaningful descriptor of choir intonation quality in the absence of individual voice separation?

Key findings

  • The proposed framework successfully models f0 dispersion as a distribution characterized by mean f0 and bandwidth, with results showing consistent trends across choir sections.
  • Bass sections exhibit the highest average f0 dispersion (26.02 cents), followed by tenors (22.22 cents), altos (22.66 cents), and sopranos (20.16 cents), matching prior ground-truth studies.
  • The framework shows robustness to variations in the number of singers per section, with no strong differences in dispersion values between 1 and 4 singers per part.
  • The method achieves consistent dispersion trends even when absolute f0 values are not perfectly accurate, indicating its utility for comparative analysis.
  • The use of spectral whitening significantly improves resolution for detecting f0 distribution peaks, especially in dense harmonic content.
  • The framework's performance is highly dependent on the accuracy of the initial multi-f0 estimation stage, suggesting that specialized voice models could further improve results.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.