[Paper Review] What's Cracking? A Review and Analysis of Deep Learning Methods for Structural Crack Segmentation, Detection and Quantification
This paper reviews deep learning methods for structural crack segmentation, detection, and quantification in civil infrastructure, analyzing supervised, semi-supervised, weakly-supervised, and unsupervised approaches. It identifies key challenges in dataset standardization, annotation quality, and evaluation metrics, and calls for unified benchmarks and temporal datasets to advance automated structural health monitoring.
Surface cracks are a very common indicator of potential structural faults. Their early detection and monitoring is an important factor in structural health monitoring. Left untreated, they can grow in size over time and require expensive repairs or maintenance. With recent advances in computer vision and deep learning algorithms, the automatic detection and segmentation of cracks for this monitoring process have become a major topic of interest. This review aims to give researchers an overview of the published work within the field of crack analysis algorithms that make use of deep learning. It outlines the various tasks that are solved through applying computer vision algorithms to surface cracks in a structural health monitoring setting and also provides in-depth reviews of recent fully, semi and unsupervised approaches that perform crack classification, detection, segmentation and quantification. Additionally, this review also highlights popular datasets used for cracks and the metrics that are used to evaluate the performance of those algorithms. Finally, potential research gaps are outlined and further research directions are provided.
Motivation & Objective
- To provide a comprehensive review of deep learning techniques applied to crack segmentation, detection, and quantification in structural health monitoring (SHM).
- To analyze the current state of supervised, semi-supervised, weakly-supervised, and unsupervised learning in crack analysis.
- To evaluate commonly used datasets and performance metrics, highlighting inconsistencies and limitations.
- To identify critical research gaps, including lack of standardized evaluation, poor annotation quality, and absence of temporal crack progression data.
- To propose future research directions, especially in weakly- and self-supervised learning and the development of longitudinal crack datasets.
Proposed method
- Systematic review of peer-reviewed literature on deep learning for crack analysis in civil infrastructure.
- Categorization of methods by learning paradigm: supervised, semi-supervised, weakly-supervised, and unsupervised.
- Analysis of performance metrics such as IoU, Dice coefficient, precision, recall, and F1-score, with critique of metric misuse and threshold dependency.
- Evaluation of popular datasets like PC-2018, CSD, and CRACK500, focusing on annotation quality and diversity.
- Examination of architectural trends, including U-Net variants, Faster R-CNN, YOLO, and transformer-based models.
- Proposal of standardized evaluation protocols and the need for temporal datasets to enable crack propagation modeling.
Experimental results
Research questions
- RQ1What are the dominant deep learning architectures and learning paradigms used in crack segmentation, detection, and quantification?
- RQ2How do current evaluation metrics affect the comparability and reproducibility of crack detection algorithms?
- RQ3What are the key limitations in existing datasets, particularly regarding annotation quality and representativeness?
- RQ4Why is the adoption of semi-, weakly-, and unsupervised learning limited in crack analysis, and how can it be improved?
- RQ5What are the major research gaps in modeling crack progression over time, and how can longitudinal datasets advance predictive maintenance?
Key findings
- Supervised deep learning methods achieve state-of-the-art performance in crack segmentation and detection, but results are difficult to compare due to inconsistent metrics and evaluation protocols.
- Many studies use loose or non-standardized metrics such as IoU with relaxed thresholds, which can lead to over-optimistic performance estimates and poor real-world generalization.
- A significant lack of standardized, high-quality datasets exists, especially for semi-, weakly-, and unsupervised learning, hindering method comparison and reproducibility.
- The absence of longitudinal datasets showing crack evolution over time limits the development of predictive models for crack propagation and proactive maintenance.
- Current annotation practices suffer from inconsistencies and low quality, which negatively impacts model performance and generalization, especially in real-world deployment.
- There is a strong need for unified evaluation standards, including fixed confidence thresholds and multi-metric reporting, to improve reproducibility and benchmarking across studies.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.