Skip to main content
QUICK REVIEW

[论文解读] Ovarian Cancer Data Analysis using Deep Learning: A Systematic Review from the Perspectives of Key Features of Data Analysis and AI Assurance

Muta Tah Hira, Mohammad A. Razzaque|arXiv (Cornell University)|Nov 20, 2023
AI in cancer detection被引用 4
一句话总结

本篇系统性综述分析了2015至2023年间96项基于深度学习(DL)的卵巢癌(OC)研究,以评估关键数据分析特征与人工智能保障(AIA)。研究发现,检测/诊断是主要焦点(71%),地理多样性有限(75%为区域特定),整合或外部验证极少(8.3%),且人工智能保障近乎完全缺失(仅2.1%的研究涉及可解释性),凸显了DL在OC研究中可泛化性、公平性与可信度方面的关键缺口。

ABSTRACT

Background and objectives: By extracting this information, Machine or Deep Learning (ML/DL)-based autonomous data analysis tools can assist clinicians and cancer researchers in discovering patterns and relationships from complex data sets. Many DL-based analyses on ovarian cancer (OC) data have recently been published. These analyses are highly diverse in various aspects of cancer (e.g., subdomain(s) and cancer type they address) and data analysis features. However, a comprehensive understanding of these analyses in terms of these features and AI assurance (AIA) is currently lacking. This systematic review aims to fill this gap by examining the existing literature and identifying important aspects of OC data analysis using DL, explicitly focusing on the key features and AI assurance perspectives. Methods: The PRISMA framework was used to conduct comprehensive searches in three journal databases. Only studies published between 2015 and 2023 in peer-reviewed journals were included in the analysis. Results: In the review, a total of 96 DL-driven analyses were examined. The findings reveal several important insights regarding DL-driven ovarian cancer data analysis: - Most studies 71% (68 out of 96) focused on detection and diagnosis, while no study addressed the prediction and prevention of OC. - The analyses were predominantly based on samples from a non-diverse population (75% (72/96 studies)), limited to a geographic location or country. - Only a small proportion of studies (only 33% (32/96)) performed integrated analyses, most of which used homogeneous data (clinical or omics). - Notably, a mere 8.3% (8/96) of the studies validated their models using external and diverse data sets, highlighting the need for enhanced model validation, and - The inclusion of AIA in cancer data analysis is in a very early stage; only 2.1% (2/96) explicitly addressed AIA through explainability.

研究动机与目标

  • 识别并分析卵巢癌(OC)研究中基于深度学习(DL)的数据分析的关键特征。
  • 评估当前基于深度学习的OC研究中人工智能保障(AIA)的现状,包括可解释性、公平性与可信度。
  • 评估OC数据分析中的多样性、整合性与验证实践,以识别关键研究缺口。
  • 通过突出预测、预防及多中心外部验证等未充分探索的领域,为未来研究提供指导。
  • 倡导全面整合人工智能保障,以确保深度学习模型在肿瘤学领域安全、公平且可靠地部署。

提出的方法

  • 采用PRISMA框架,在三个经过同行评审的期刊数据库中开展系统性文献综述。
  • 筛选2015至2023年间发表的96项应用深度学习于卵巢癌数据分析的研究。
  • 根据数据模态(如影像、组学、临床数据)、分析目标(如诊断、预后)与方法学特征对研究进行分类。
  • 评估模型验证实践,包括内部与外部验证,以及对多样化人口群体的使用情况。
  • 评估人工智能保障(AIA)组件的纳入情况,如可解释性、公平性、安全性与可信度。
  • 将研究发现整合为关于数据整合、模型可泛化性及OC DL应用中研究缺口的主题性洞察。

实验结果

研究问题

  • RQ12015至2023年间,基于深度学习的卵巢癌研究中,主导的数据模态与分析目标是什么?
  • RQ2在OC研究中,DL模型在多大程度上使用外部、多样化或多中心数据集进行了验证?
  • RQ3当前基于深度学习的OC研究在多大程度上全面涵盖了人工智能保障(AIA)原则,如可解释性、公平性与安全性?
  • RQ4有多少比例的研究采用了整合的多组学或混合DL-ML方法?这些方法在性能与可泛化性方面如何比较?
  • RQ5在DL驱动的卵巢癌数据分析中,关键的方法论与伦理缺口是什么,特别是在人群多样性与模型鲁棒性方面?

主要发现

  • 71%(96项中的68项)的研究仅专注于检测与诊断,无一项研究涉及卵巢癌的预测或预防。
  • 75%(96项中的72项)的研究使用单一地理区域或国家的数据,引发对人群偏差与泛化能力有限的担忧。
  • 仅33%(96项中的32项)的研究进行了整合性分析,且大多数局限于同质数据类型,如临床或组学数据。
  • 仅8.3%(96项中的8项)的研究使用外部且多样化数据集验证其模型,表明在鲁棒性与实际应用场景适配性方面存在重大缺口。
  • 人工智能保障(AIA)仍基本未被关注,仅有2.1%(96项中的2项)的研究明确纳入可解释性,且无任何研究涉及公平性或安全性等其他AIA维度。
  • 本综述指出,未来研究亟需优先关注多中心、多样化人群的验证,整合异质性数据的分析方法,以及在卵巢癌DL应用中建立全面的人工智能保障框架。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。