[论文解读] Measuring the Quality of Explanations: The System Causability Scale (SCS). Comparing Human and Machine Explanations
本文提出了系统可理解性量表(SCS),一种基于系统可用性量表(SUS)的10项李克特量表工具,用于评估人机交互中AI解释的质量,特别是在医疗AI领域。该量表表现出较强的信度(Cronbach’s alpha = .91)和效度,通过在弗朗西斯科风险评估工具(Framingham Risk Tool)上的应用,获得了0.86的高SCS得分,表明解释的可理解性感知水平较高。
Recent success in Artificial Intelligence (AI) and Machine Learning (ML) allow problem solving automatically without any human intervention. Autonomous approaches can be very convenient. However, in certain domains, e.g., in the medical domain, it is necessary to enable a domain expert to understand, <i>why</i> an algorithm came up with a certain result. Consequently, the field of Explainable AI (xAI) rapidly gained interest worldwide in various domains, particularly in medicine. Explainable AI studies transparency and traceability of opaque AI/ML and there are already a huge variety of methods. For example with layer-wise relevance propagation relevant parts of inputs to, and representations in, a neural network which caused a result, can be highlighted. This is a first important step to ensure that end users, e.g., medical professionals, assume responsibility for decision making with AI/ML and of interest to professionals and regulators. Interactive ML adds the component of human expertise to AI/ML processes by enabling them to re-enact and retrace AI/ML results, e.g. let them check it for plausibility. This requires new human-AI interfaces for explainable AI. In order to build effective and efficient interactive human-AI interfaces we have to deal with the question of <i>how to evaluate the quality of explanations</i> given by an explainable AI system. In this paper we introduce our System Causability Scale to measure the quality of explanations. It is based on our notion of Causability (Holzinger et al. in Wiley Interdiscip Rev Data Min Knowl Discov 9(4), 2019) combined with concepts adapted from a widely-accepted usability scale.
研究动机与目标
- 为解决可解释AI(xAI)系统中解释质量评估缺乏标准化工具的问题。
- 开发一种可靠、快速且心理测量学上有效的工具,用于衡量领域专家对AI解释的因果可理解性与可用性。
- 通过评估解释在支持理解、信任和决策方面的作用,支持高风险领域(如医学)中有效人机界面的设计。
- 将SCS与既有的可用性度量进行对比验证,并展示其在真实世界医疗预测模型中的适用性。
提出的方法
- 借鉴系统可用性量表(SUS)的框架,开发了一种聚焦于可理解性而非一般可用性的10项李克特量表工具。
- 设计题目以评估关键维度:因果因素的完整性、上下文中的清晰度、细节调整的灵活性、对额外支持的独立性,以及解释提供的及时性。
- 将SCS应用于广泛使用的临床预测模型——弗朗西斯科风险评估工具(FRT),以在真实医疗情境中评估解释质量。
- 使用Cronbach’s alpha计算内部一致性,并通过SCS与原始SUS的相关性来评估收敛效度。
- 每项使用5点李克特量表(1 = 强烈不同意 至 5 = 强烈同意),总分标准化为0–1量表。
- 通过一名在职医生的试点评估,检验实际应用场景中解释的可用性和感知质量。
实验结果
研究问题
- RQ1如何系统性地测量人机交互中解释的质量,以体现其因果透明度和可用性?
- RQ2SCS与SUS等既有的可用性度量的相关性在多大程度上支持其效度?
- RQ3SCS能否有效区分真实世界医疗AI系统(如弗朗西斯科风险评估工具)中解释的感知质量?
- RQ4SCS在不同用户和情境下测量解释质量的可靠性如何?
主要发现
- 系统可理解性量表(SCS)表现出极高的内部一致性,Cronbach’s alpha为.91,表明其信度很强。
- SCS与原始系统可用性量表(SUS)之间存在极高的相关性(r = .985),支持其收敛效度。
- 在弗朗西斯科风险评估工具上的应用中,SCS获得了0.86的标准化得分,表明解释的可理解性感知水平很高。
- 如‘在上下文中可理解’(评分5)和‘无需外部支持’(评分5)等项目得分最高,反映出解释的清晰度和独立性较强。
- SCS能够识别出‘使用知识库支持’(评分3)和‘快速学习’(评分3)等维度的较弱表现,突显了界面改进的潜在方向。
- SCS与第二个派生量表之间存在中等程度的相关性(r = .664),表明其测量的是与一般可用性相关但又有所区别的构念。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。