[Paper Review] Anomalous Sound Detection with Machine Learning: A Systematic Review
This systematic review synthesizes 31 studies on machine learning for anomalous sound detection (ASD) from 2010–2020, identifying key datasets, feature extraction methods, ML models, and evaluation metrics. It finds that MFCCs, autoencoders, and CNNs are the most prevalent techniques, with AUC and F1-score as standard evaluation measures, offering a comprehensive state-of-the-art overview for researchers in audio anomaly detection.
Anomalous sound detection (ASD) is the task of identifying whether the sound emitted from an object is normal or anomalous. In some cases, early detection of this anomaly can prevent several problems. This article presents a Systematic Review (SR) about studies related to Anamolous Sound Detection using Machine Learning (ML) techniques. This SR was conducted through a selection of 31 (accepted studies) studies published in journals and conferences between 2010 and 2020. The state of the art was addressed, collecting data sets, methods for extracting features in audio, ML models, and evaluation methods used for ASD. The results showed that the ToyADMOS, MIMII, and Mivia datasets, the Mel-frequency cepstral coefficients (MFCC) method for extracting features, the Autoencoder (AE) and Convolutional Neural Network (CNN) models of ML, the AUC and F1-score evaluation methods were most cited.
Motivation & Objective
- To map the current state of the art in machine learning for anomalous sound detection (ASD) between 2010 and 2020.
- To identify the most frequently used datasets, feature extraction methods, machine learning models, and evaluation metrics in ASD research.
- To provide a structured, evidence-based overview to guide future research and development in audio-based anomaly detection.
- To support the design of more effective and standardized ASD systems by synthesizing trends and gaps in existing literature.
Proposed method
- Conducted a systematic literature review (SLR) following PRISMA guidelines across multiple academic databases including IEEE Xplore, ACM, Scopus, and DCASE.
- Defined inclusion and exclusion criteria: studies must focus on ASD using ML, be published in English between 2010 and 2020, and include full-text access.
- Collected data on ML categories, anomaly detection types (supervised, semi-supervised, unsupervised), datasets, feature extraction techniques, models, and evaluation metrics.
- Extracted and analyzed data from 31 primary studies after screening titles, abstracts, and full texts for relevance and methodological quality.
- Used standardized data extraction and quality assessment forms to ensure consistency and reliability across studies.
- Categorized and summarized findings by dataset, feature extraction, model architecture, and evaluation performance to identify dominant trends.
Experimental results
Research questions
- RQ1Which machine learning models are most commonly used in anomalous sound detection research between 2010 and 2020?
- RQ2Which audio feature extraction methods are most prevalent in ASD studies, and how do they compare in performance?
- RQ3Which datasets are most frequently used in ASD research, and what are their characteristics in terms of domain and labeling?
- RQ4What evaluation metrics are standard in ASD research, and how do they reflect model generalization and robustness?
- RQ5How has the use of supervised, semi-supervised, and unsupervised learning evolved in ASD over the past decade?
Key findings
- The ToyADMOS, MIMII, and Mivia datasets were the most frequently cited in ASD research, indicating their prominence in benchmarking.
- Mel-frequency cepstral coefficients (MFCCs) were the most widely used feature extraction method across the reviewed studies.
- Autoencoders (AE) and Convolutional Neural Networks (CNNs) were the most prevalent machine learning models, especially in unsupervised and semi-supervised settings.
- The Area Under the Curve (AUC) and F1-score were the most commonly reported evaluation metrics, reflecting their importance in assessing model performance.
- A majority of studies employed semi-supervised or unsupervised learning, highlighting the challenge of obtaining large-scale labeled anomaly data.
- There was a notable trend toward self-supervised and weakly supervised approaches, particularly in industrial and real-world audio applications.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.