Skip to main content
QUICK REVIEW

[Paper Review] On the Characterization of Expressive Performance in Classical Music: First Results of the Con Espressione Game

Carlos Cancino-Chacón, Silvan Peter|arXiv (Cornell University)|Aug 5, 2020
Music and Audio Processing4 citations
TL;DR

This paper introduces the Con Espressione Game (CEG), an online dataset of 1,500 listener-generated adjectives describing expressive character in 45 classical piano performances. Using natural language descriptions, it identifies four key expressive dimensions—particularly 'calm vs. agitated'—and shows that expressive parameters (e.g., dynamics, tempo) significantly predict these dimensions, especially for the most interpretable axis.

ABSTRACT

A piece of music can be expressively performed, or interpreted, in a variety of ways. With the help of an online questionnaire, the Con Espressione Game, we collected some 1,500 descriptions of expressive character relating to 45 performances of 9 excerpts from classical piano pieces, played by different famous pianists. More specifically, listeners were asked to describe, using freely chosen words (preferably: adjectives), how they perceive the expressive character of the different performances. In this paper, we offer a first account of this new data resource for expressive performance research, and provide an exploratory analysis, addressing three main questions: (1) how similarly do different listeners describe a performance of a piece? (2) what are the main dimensions (or axes) for expressive character emerging from this?; and (3) how do measurable parameters of a performance (e.g., tempo, dynamics) and mid- and high-level features that can be predicted by machine learning models (e.g., articulation, arousal) relate to these expressive dimensions? The dataset that we publish along with this paper was enriched by adding hand-corrected score-to-performance alignments, as well as descriptive audio features such as tempo and dynamics curves.

Motivation & Objective

  • To investigate inter-listener consistency in describing expressive character of classical piano performances using free-text adjectives.
  • To identify the main perceptual dimensions underlying listeners' descriptions of expressive performance.
  • To examine the relationship between measurable audio performance parameters (e.g., tempo, dynamics) and mid- and high-level features with perceived expressive character.
  • To create and release a richly annotated dataset enriched with score-to-performance alignments and descriptive audio features for future research in expressive performance.

Proposed method

  • Collected 1,500 free-text adjective descriptions from listeners via the online Con Espressione Game (CEG) for 45 performances of 9 classical piano excerpts.
  • Applied multidimensional scaling (MDS) to reduce the semantic space of descriptors into four primary expressive character dimensions.
  • Used multiple linear regression (MLR) with Zheng-Loh variable selection to test predictive power of three feature sets: expressive parameters, mid-level features (e.g., articulation, arousal), and high-level features (e.g., valence, energy).
  • Enriched the dataset with hand-corrected score-to-performance alignments and audio features such as tempo and dynamics curves.
  • Employed NLP techniques to assess semantic similarity, though noted limitations due to metaphor-laden and context-dependent terms.
  • Analyzed listener preferences and training effects to explore how musical expertise influences description complexity and perception.

Experimental results

Research questions

  • RQ1To what extent do listeners agree in their descriptive characterizations of the same expressive performance?
  • RQ2What are the primary perceptual dimensions that organize listeners’ free-text descriptions of expressive character?
  • RQ3How do measurable performance parameters (e.g., tempo, dynamics) and mid- and high-level features (e.g., arousal, valence) relate to the identified expressive dimensions?
  • RQ4How do listener preferences and musical training influence the complexity and content of descriptive responses?

Key findings

  • Listeners show moderate to high consistency in describing the expressive character of the same performance, with a small positive correlation between description complexity and musical training.
  • Four main expressive character dimensions were identified: (1) 'gentle/calm' vs. 'hectic/agitated', (2) 'warm' vs. 'cold', (3) 'flowing' vs. 'rigid', and (4) 'elegant' vs. 'clumsy'.
  • Expressive parameters (e.g., dynamics, tempo) significantly predict all four dimensions, with medium effect sizes (R²), and are most strongly linked to Dimension 1 ('calm vs. agitated').
  • Mid-level features predict Dimensions 1 and 4; high-level features predict Dimensions 1 and 3, indicating that Dimension 1 is most systematically related to measurable performance features.
  • Loudness outliers and irregular valence curves are associated with perceptions of 'agitated' or 'irregular' character, while softer, stable dynamics correlate with 'calm' or 'graceful' descriptions.
  • Deadpan performances and Glenn Gould’s idiosyncratic interpretations were least preferred, suggesting listener preference for quantitatively average expressive styles.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.