[Paper Review] HEAR: Holistic Evaluation of Audio Representations
HEAR benchmarks general-purpose audio embeddings by evaluating 29 models across 19 diverse tasks spanning speech, environmental sound, and music, using a common API and frozen embeddings for downstream classifiers.
What audio embedding approach generalizes best to a wide range of downstream tasks across a variety of everyday domains without fine-tuning? The aim of the HEAR benchmark is to develop a general-purpose audio representation that provides a strong basis for learning in a wide variety of tasks and scenarios. HEAR evaluates audio representations using a benchmark suite across a variety of domains, including speech, environmental sound, and music. HEAR was launched as a NeurIPS 2021 shared challenge. In the spirit of shared exchange, each participant submitted an audio embedding model following a common API that is general-purpose, open-source, and freely available to use. Twenty-nine models by thirteen external teams were evaluated on nineteen diverse downstream tasks derived from sixteen datasets. Open evaluation code, submitted models and datasets are key contributions, enabling comprehensive and reproducible evaluation, as well as previously impossible longitudinal studies. It still remains an open question whether one single general-purpose audio representation can perform as holistically as the human ear.
Motivation & Objective
- Motivate the need for a single general-purpose audio representation that transfers across diverse tasks without fine-tuning.
- Provide a broad, open, reproducible evaluation framework for audio representations across multiple domains.
- Enable rapid iteration by offering a common API, open datasets, and evaluation code for cross-model comparison.
Proposed method
- Provide a common HEAR API to wrap embeddings from diverse models for downstream tasks.
- Evaluate embeddings using two task types: scene-based classification and timestamp-based sound event detection, with frozen embeddings fed to a shallow MLP.
- Preprocess datasets to a common format with standard splits and open licenses to facilitate reproducibility.
- Use open-source evaluation code and standardized preprocessing to enable cross-model comparisons and longitudinal studies.
Experimental results
Research questions
- RQ1Can a single general-purpose audio representation perform holistically across speech, environmental sounds, and music tasks?
- RQ2How do different pretraining regimes (supervised vs self-supervised, single-domain vs multimodal) influence cross-task transferability?
- RQ3Which layers or fusion strategies of embeddings yield the best generalization without fine-tuning?
Key findings
- Several models benefit from non-final or fused layers, not just the last layer, for transfer to diverse tasks.
- Models incorporating pitch-specific embeddings (e.g., CREPE) perform well on pitch-related tasks like NSynth Pitch and Maestro.
- AudioSet-pretrained semantic-object tagging models correlate with performance on broad-domain tasks like ESC-50 and GTZAN genre tagging.
- CP-JKU PaSST models achieved state-of-the-art mean average precision on FSD50K without fine-tuning.
- Strong speech models often excel on speech-related tasks such as emotion recognition and language identification.
- The benchmark emphasizes openness, reproducibility, and the challenge of achieving true holistic audio representations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.