Skip to main content
QUICK REVIEW

[Paper Review] Maximum entropy models for generation of expressive music

Simon Moulieras, François Pachet|arXiv (Cornell University)|Oct 12, 2016
Music Technology and Sound Studies10 references3 citations
TL;DR

This paper proposes a Maximum Entropy (MaxEnt) model to generate expressive music by learning microtiming deviations, duration shifts, loudness changes, and metrical positions from professional pianist performances of jazz, pop, and Latin jazz. The model captures local correlations between neighboring notes and generates musically expressive performances that are statistically indistinguishable from human performances in listening tests, with participants significantly preferring MaxEnt-generated melodies over non-expressive or randomly expressive versions.

ABSTRACT

In the context of contemporary monophonic music, expression can be seen as the difference between a musical performance and its symbolic representation, i.e. a musical score. In this paper, we show how Maximum Entropy (MaxEnt) models can be used to generate musical expression in order to mimic a human performance. As a training corpus, we had a professional pianist play about 150 melodies of jazz, pop, and latin jazz. The results show a good predictive power, validating the choice of our model. Additionally, we set up a listening test whose results reveal that on average, people significantly prefer the melodies generated by the MaxEnt model than the ones without any expression, or with fully random expression. Furthermore, in some cases, MaxEnt melodies are almost as popular as the human performed ones.

Motivation & Objective

  • To model musical expression as deviations from symbolic scores using a principled probabilistic framework.
  • To capture local, translation-invariant correlations between note-level expressive parameters (timing, duration, loudness, metrical position).
  • To develop a generative model that mimics human performance without relying on global structural patterns.
  • To evaluate the perceptual quality of generated music through controlled listening tests.
  • To demonstrate that MaxEnt models can produce expressive music that is close to human performance in listener preference.

Proposed method

  • Apply the Principle of Maximum Entropy to construct a probability distribution that maximizes entropy under observed constraints from performance data.
  • Model discrete variables (e.g., metrical position) and continuous variables (e.g., onset deviation, duration deviation, loudness) using separate MaxEnt distributions with exponential family forms.
  • Use local, translation-invariant features—where each note’s expressive parameters depend only on its immediate neighbors—thereby enforcing a local texture assumption.
  • Define sufficient statistics (observables) such as onset deviation, duration deviation, loudness, and metrical position to capture relevant expressive features.
  • Estimate model parameters via maximum likelihood using a corpus of 150 professionally performed melodies across jazz, pop, and Latin jazz genres.
  • Generate new performances by sampling from the learned MaxEnt distribution, preserving statistical dependencies between neighboring notes.

Experimental results

Research questions

  • RQ1Can a Maximum Entropy model effectively capture and generate expressive musical performances from real-world data?
  • RQ2To what extent do local correlations between neighboring notes account for expressive timing and dynamics in contemporary music?
  • RQ3How perceptually similar are MaxEnt-generated performances to human performances, especially compared to non-expressive or randomly expressive renditions?
  • RQ4Can the model generalize across diverse musical styles such as jazz, pop, and Latin jazz?
  • RQ5What is the relative preference of listeners for MaxEnt-generated music versus human performances?

Key findings

  • The MaxEnt model demonstrated strong predictive power on the training data, validating its ability to capture salient expressive patterns.
  • In a listening test, participants significantly preferred MaxEnt-generated melodies over both non-expressive and fully random expressive versions.
  • In some cases, MaxEnt-generated melodies were almost as preferred as the original human performances, indicating high perceptual plausibility.
  • The model successfully captured local expressive correlations without requiring long-range structural assumptions, supporting the hypothesis that expression is primarily a local texture.
  • The use of continuous variables (onset, duration, loudness) within the MaxEnt framework allowed for accurate modeling of microtiming deviations in real-world performances.
  • The translation-invariant assumption—where expressive features depend only on local context—proved effective and sufficient for generating convincing musical expression.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.