Skip to main content
QUICK REVIEW

[Paper Review] The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use

Bob L. Sturm|VBN Forskningsportal (Aalborg Universitet)|Jun 6, 2013
Music and Audio ProcessingComputer Science131 references86 citations
TL;DR

This paper critically evaluates the GTZAN dataset—widely used for music genre recognition (MGR)—by identifying its faults (repetitions, mislabelings, distortions), demonstrating that these flaws invalidate direct performance comparisons across MGR systems, and showing that even state-of-the-art systems perform inconsistently when evaluated on GTZAN. The study concludes that GTZAN should not be abandoned but used with full awareness of its content and flaws to ensure valid evaluation in music machine listening research.

ABSTRACT

The GTZAN dataset appears in at least 100 published works, and is the most-used public dataset for evaluation in machine listening research for music genre recognition (MGR). Our recent work, however, shows GTZAN has several faults (repetitions, mislabelings, and distortions), which challenge the interpretability of any result derived using it. In this article, we disprove the claims that all MGR systems are affected in the same ways by these faults, and that the performances of MGR systems in GTZAN are still meaningfully comparable since they all face the same faults. We identify and analyze the contents of GTZAN, and provide a catalog of its faults. We review how GTZAN has been used in MGR research, and find few indications that its faults have been known and considered. Finally, we rigorously study the effects of its faults on evaluating five different MGR systems. The lesson is not to banish GTZAN, but to use it with consideration of its contents.

Motivation & Objective

  • To identify and catalog the faults in the GTZAN dataset, including repetitions, mislabelings, and audio distortions.
  • To challenge the widely held assumption that all MGR systems are equally affected by GTZAN's flaws, undermining the validity of performance comparisons.
  • To analyze how these faults impact the evaluation of five distinct MGR systems, revealing inconsistent and misleading performance rankings.
  • To provide a comprehensive metadata catalog for GTZAN excerpts to improve transparency and reproducibility in future research.
  • To advocate for responsible use of GTZAN in MGR and related tasks by emphasizing content-aware evaluation design over blind reliance on benchmark scores.

Proposed method

  • Conducted a systematic analysis of all 1,000 audio excerpts in GTZAN to identify repetitions, mislabelings, and audio distortions through audio signal processing and manual listening validation.
  • Extended prior metadata work by creating detailed metadata for 110 additional excerpts, enabling accurate content-based evaluation of genre labels.
  • Formally defined and classified mislabelings using audio content analysis and expert listening, distinguishing between true genre misalignments and perceptual ambiguities.
  • Evaluated five state-of-the-art MGR systems (including MAPsCAT and SRCAM) on GTZAN under controlled conditions to measure the impact of dataset faults on classification accuracy.
  • Established upper bounds for classification performance under ideal conditions by analyzing the most consistent and correctly labeled excerpts.
  • Proposed a framework for future evaluation that prioritizes content-aware experimental design over reliance on aggregate metrics in flawed datasets.

Experimental results

Research questions

  • RQ1To what extent do repetitions, mislabelings, and distortions in GTZAN invalidate performance comparisons across MGR systems?
  • RQ2Do all MGR systems respond to GTZAN's faults in the same way, or do some benefit or suffer disproportionately?
  • RQ3How do the identified faults affect the classification accuracy of diverse MGR systems, and can we quantify the performance degradation or inflation?
  • RQ4Can the GTZAN dataset still be useful for future research if its flaws are acknowledged and addressed in evaluation design?
  • RQ5What are the upper bounds of performance on GTZAN for a 'perfect' MGR system, and how do they compare to reported results in the literature?

Key findings

  • The study disproves the claim that all MGR systems are equally affected by GTZAN’s faults, showing that performance rankings are unreliable and not meaningfully comparable.
  • SRCAM and MAPsCAT—previously reported to achieve 83% accuracy—ranked at the bottom of the performance spectrum after accounting for dataset faults, indicating that prior results were inflated or misleading.
  • Over 100 published works used GTZAN for MGR evaluation, but only five acknowledged any awareness of its content issues, and none systematically considered the musical content in evaluation.
  • The dataset contains 110 previously unidentified excerpts with mislabelings or distortions, and several tracks are repeated or incorrectly categorized (e.g., classical pieces mislabeled as jazz or rock).
  • The upper bound for a perfect MGR system on GTZAN is estimated to be below 90% accuracy due to inherent ambiguities and inconsistencies in the data.
  • The paper confirms that dataset size alone does not resolve fundamental issues—large datasets can still contain uncontrolled variables, and GTZAN’s flaws are representative of real-world data challenges in music machine listening.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.