[论文解读] Explainable AI for clinical risk prediction: a survey of concepts, methods, and modalities
本综述提出了一套全面的可解释人工智能(XAI)框架,用于临床风险预测,整合了多模态医疗数据中的可解释性、公平性和透明度。该框架倡导使用合成数据集进行端到端验证、多方法可解释性分析以及开放科学实践,以增强信任度、可靠性和临床应用。
Recent advancements in AI applications to healthcare have shown incredible promise in surpassing human performance in diagnosis and disease prognosis. With the increasing complexity of AI models, however, concerns regarding their opacity, potential biases, and the need for interpretability. To ensure trust and reliability in AI systems, especially in clinical risk prediction models, explainability becomes crucial. Explainability is usually referred to as an AI system's ability to provide a robust interpretation of its decision-making logic or the decisions themselves to human stakeholders. In clinical risk prediction, other aspects of explainability like fairness, bias, trust, and transparency also represent important concepts beyond just interpretability. In this review, we address the relationship between these concepts as they are often used together or interchangeably. This review also discusses recent progress in developing explainable models for clinical risk prediction, highlighting the importance of quantitative and clinical evaluation and validation across multiple common modalities in clinical practice. It emphasizes the need for external validation and the combination of diverse interpretability methods to enhance trust and fairness. Adopting rigorous testing, such as using synthetic datasets with known generative factors, can further improve the reliability of explainability methods. Open access and code-sharing resources are essential for transparency and reproducibility, enabling the growth and trustworthiness of explainable research. While challenges exist, an end-to-end approach to explainability in clinical risk prediction, incorporating stakeholders from clinicians to developers, is essential for success.
研究动机与目标
- 为应对在临床风险预测中使用复杂人工智能模型时日益增长的可解释性需求,此类决策直接影响患者结果。
- 阐明临床人工智能系统中可解释性、可理解性、公平性、偏见与透明度之间的相互作用。
- 通过定量评估和临床评估,对电子健康记录(EHRs)、医学影像和文本等多样化临床模态中的XAI方法进行评估与验证。
- 倡导外部验证、合成数据集测试以及代码共享,以提高XAI研究的可靠性与可重现性。
- 推动端到端、利益相关者广泛参与的方法——涵盖临床医生、患者及开发人员——以确保信任与实际应用中的可用性。
提出的方法
- 系统性回顾临床人工智能中可解释性、可理解性、公平性与透明度的概念,明确区分其相互关系。
- 分析一系列XAI技术,包括事后解释方法(如LIME、SHAP)、内在可解释模型(如XGBoost、决策树)以及基于规则的系统。
- 提出使用具有已知生成因素的合成数据集,以在临床部署前严格测试XAI方法的可靠性。
- 整合定量评估指标,如AUROC和F1分数,并针对特征排序和基于破坏的鲁棒性测试进行适应性调整。
- 强调需要采用多方法可解释性——结合多种XAI技术——以提升鲁棒性并减少偏差。
- 呼吁开放获取与代码共享,以确保XAI研究的透明度、可重现性与可信度。

实验结果
研究问题
- RQ1如何在不同数据模态中系统评估可解释人工智能方法在临床风险预测中的可靠性与有效性?
- RQ2在人工智能驱动的临床决策中,可理解性、公平性、透明度与信任之间存在何种关系?
- RQ3在已知潜在生成因素的前提下,合成数据集在医疗XAI方法验证中可发挥多大作用?
- RQ4如何使临床医生和患者在可解释人工智能系统的设计与评估中实现有意义的参与?
- RQ5监管框架与开放科学实践在推动医疗领域可信且可重现的XAI发展中发挥何种作用?
主要发现
- 可解释人工智能对于建立临床风险预测模型的信任至关重要,尤其是在模型错误或偏见可能带来高风险后果的背景下。
- 仅依赖事后解释方法是不够的;必须结合多种可理解性技术,才能实现对模型的全面理解。
- 在具有已知生成因素的合成数据集上测试XAI方法,可显著提升其可靠性与验证信心。
- 采用损坏特征与排序指标(如AUROC、F1)的定量评估框架,对特征重要性解释具有良好的评估效果。
- 开放获取与代码共享对于XAI研究的可重现性以及长期可信度至关重要,尤其是在医疗等敏感领域。
- 尽管已取得进展,但临床人工智能中许多可解释性声明仍被夸大,若缺乏严格的验证与利益相关者参与,其在真实临床环境中的采纳仍十分有限。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。