Skip to main content
QUICK REVIEW

[Paper Review] Deep Predictive Models in Interactive Music

Charles Martín, Kai Olav Ellefsen|arXiv (Cornell University)|Jan 31, 2018
Music Technology and Sound Studies13 references3 citations
TL;DR

This paper investigates how deep predictive models in digital musical instruments (DMIs) extend human cognitive prediction in musical performance, enabling real-time musical continuation, adaptive control, and ensemble interaction. By framing mapping and modelling as predictive tasks, the authors demonstrate that recurrent neural networks (RNNs) and other machine learning models enhance expressive, interactive music systems with long-term memory and intelligent responsiveness.

ABSTRACT

Musical performance requires prediction to operate instruments, to perform in groups and to improvise. In this paper, we investigate how a number of digital musical instruments (DMIs), including two of our own, have applied predictive machine learning models that assist users by predicting unknown states of musical processes. We characterise these predictions as focussed within a musical instrument, at the level of individual performers, and between members of an ensemble. These models can connect to existing frameworks for DMI design and have parallels in the cognitive predictions of human musicians. We discuss how recent advances in deep learning highlight the role of prediction in DMIs, by allowing data-driven predictive models with a long memory of past states. The systems we review are used to motivate musical use-cases where prediction is a necessary component, and to highlight a number of challenges for DMI designers seeking to apply deep predictive models in interactive music systems of the future.

Motivation & Objective

  • To investigate how predictive machine learning models support musical performance in digital musical instruments (DMIs).
  • To identify and categorize predictive models used in DMIs across instrument, performer, and ensemble levels.
  • To frame mapping and modelling in interactive music as forms of prediction, linking them to cognitive processes in musicians.
  • To highlight challenges and opportunities in integrating deep learning for future DMI design.
  • To advocate for prediction as a unifying framework in interactive music systems, especially with end-to-end learning from sensor input to sound output.

Proposed method

  • Categorize existing DMI systems by level of prediction: instrument-level (e.g., gesture-to-sound mapping), performer-level (e.g., continuation of musical phrases), and ensemble-level (e.g., predicting group interactions).
  • Review 18 DMI systems from the literature and two in-house systems (GloveTalk II, PiaF), analyzing their machine learning models, inputs, and outputs.
  • Use recurrent neural networks (RNNs), support vector machines (SVMs), hidden Markov models (HMMs), and other models to model musical sequences and map sensor data to sound parameters.
  • Frame musical mapping (control-to-sound) and modelling (representation of musical processes) as distinct but related forms of prediction.
  • Analyze how models handle temporal structures such as rhythm, harmony, and melody through long-term memory and context-aware processing.
  • Propose that deep learning models, especially RNNs, enable end-to-end predictive systems where sensor data directly drives expressive, adaptive sound generation.

Experimental results

Research questions

  • RQ1How do deep predictive models in DMIs support and extend cognitive prediction in musical performance?
  • RQ2In what ways can mapping and modelling in interactive music be reinterpreted as predictive tasks?
  • RQ3What are the key challenges in designing DMIs that use predictive models for real-time, expressive musical interaction?
  • RQ4How do different machine learning models (e.g., RNNs, HMMs, SVMs) support prediction at instrument, performer, and ensemble levels?
  • RQ5What design principles and frameworks are needed to integrate predictive intelligence into future interactive music systems?

Key findings

  • Predictive models in DMIs, especially RNNs, enable long-term memory of musical states, allowing for context-aware musical continuation and adaptation.
  • Systems like AI Duet and RoboJam use RNNs to continue musical phrases in real time, demonstrating effective performer-level prediction with symbolic music input.
  • Ensemble-level prediction is supported by models such as MalLo (predicting percussion strokes) and Neural Touchscreen Ensemble (predicting gesture classes), showing feasibility in collaborative improvisation.
  • The use of fNIRS and BCI in BRAAHMS demonstrates that brain signals can be used as input for predictive harmonic accompaniment, expanding sensor modalities.
  • Despite the potential, deep learning models like RNNs are not yet widely adopted in DMIs, indicating a gap in practical implementation.
  • The integration of predictive models enhances musical expressivity and enables new creative affordances, particularly in live performance and improvisation settings.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.