[论文解读] A survey, review, and future trends of skin lesion segmentation and classification
本篇全面综述分析了2011年至2022年间594项关于皮肤病变分割与分类的研究,评估了数据集、预处理、深度学习架构、损失函数、超参数及评估指标。研究发现ISIC数据集占主导地位,深度学习在传统机器学习之上占据主导地位,并建议整合临床部署与可解释人工智能(XAI),以提升实际应用采纳率和皮肤科医生的信任度。
The Computer-aided Diagnosis or Detection (CAD) approach for skin lesion analysis is an emerging field of research that has the potential to alleviate the burden and cost of skin cancer screening. Researchers have recently indicated increasing interest in developing such CAD systems, with the intention of providing a user-friendly tool to dermatologists to reduce the challenges encountered or associated with manual inspection. This article aims to provide a comprehensive literature survey and review of a total of 594 publications (356 for skin lesion segmentation and 238 for skin lesion classification) published between 2011 and 2022. These articles are analyzed and summarized in a number of different ways to contribute vital information regarding the methods for the development of CAD systems. These ways include relevant and essential definitions and theories, input data (dataset utilization, preprocessing, augmentations, and fixing imbalance problems), method configuration (techniques, architectures, module frameworks, and losses), training tactics (hyperparameter settings), and evaluation criteria. We intend to investigate a variety of performance-enhancing approaches, including ensemble and post-processing. We also discuss these dimensions to reveal their current trends based on utilization frequencies. In addition, we highlight the primary difficulties associated with evaluating skin lesion segmentation and classification systems using minimal datasets, as well as the potential solutions to these difficulties. Findings, recommendations, and trends are disclosed to inform future research on developing an automated and robust CAD system for skin lesion analysis.
研究动机与目标
- 提供2011年至2022年间594篇关于皮肤病变分割(SLS)与分类(SLC)研究的系统性综述。
- 识别皮肤病变分析中数据处理、模型架构、训练策略和评估基准的关键趋势。
- 突出现有研究在临床转化和可解释人工智能(XAI)整合方面的空白。
- 通过推荐数据集使用、模型优化和评估框架的最佳实践,为未来研究提供指导。
- 倡导开发基于稳健、可解释人工智能模型的用户友好型、实时临床应用。
提出的方法
- 基于Google Scholar,对2011年至2022年间594项研究(其中356项针对SLS,238项针对SLC)进行系统性文献综述。
- 根据输入数据(如ISIC数据集)、预处理、数据增强和类别不平衡处理技术,对方法进行分类与分析。
- 对深度学习架构(如FCN、U-Net、ResNet、SENet)、损失函数(如DSC、Jaccard)和优化器(如Adam、SGD)进行分类。
- 评估超参数设置,包括学习率、批量大小和训练周期数(通常为100–200)。
- 评估评估指标(如DSC、AUC、F1-score、MCC)以及定性、定量和TCT(技术、临床和技术)评估框架的使用情况。
- 分析XAI技术(如Grad-CAM、LIME、SHAP)及其在临床决策支持和信任评估中整合程度有限的问题。
实验结果
研究问题
- RQ12011年至2022年间,皮肤病变分割与分类研究中最常见的数据集、预处理技术和数据增强策略是什么?
- RQ2在SLS和SLC任务中,深度学习架构与损失函数在性能和采用率方面如何演变?
- RQ3最常使用的评估指标有哪些?它们在反映模型鲁棒性和临床相关性方面表现如何?
- RQ4为何在594项综述研究中缺乏真实临床环境的部署?哪些障碍阻碍了向临床实践的转化?
- RQ5可解释人工智能(XAI)方法在SLS和SLC流程中的整合程度如何?它们对皮肤科医生信任度和模型可解释性有何影响?
主要发现
- ISIC数据集是最广泛使用的研究基准,出现在绝大多数综述研究中,且被视为当前SLS与SLC研究的代表性数据集。
- 深度学习模型,特别是基于U-Net和ResNet的架构,配合Dice损失(DSC)或Jaccard指数(JI),已成为主流方法,因其卓越性能和端到端学习能力而被广泛采用。
- 自适应优化器(如Adam)和采用学习率调度的SGD最为常用,训练周期数通常在100至200之间。
- DSC、AUC和F1-score等评估指标最常被报告,但直接定量评估、定性评估以及TCT(技术、临床和技术)评估框架在实践中仍使用不足。
- 仅有少数研究整合了XAI技术,且更少研究评估其对诊断准确率或皮肤科医生接受度的影响,表明存在关键研究空白。
- 尽管研究数量庞大,但594篇综述论文中无一开发出可临床部署的、用户友好的应用,凸显了研究与实际临床实施之间的重大脱节。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。