[论文解读] Human-Centered Tools for Coping with Imperfect Algorithms during Medical Decision-Making
本文介紹了以人為中心的工具,使病理學家能在醫學影像檢索過程中修復和調整不完善的機器學習演算法,從而提高診斷相關性和信任度。經病理學家評估,這些工具在不損失準確性的前提下,提升了診斷實用性與使用者信心,並使使用者能即時診斷演算法錯誤、探索模型行為。
Machine learning (ML) is increasingly being used in image retrieval systems for medical decision making. One application of ML is to retrieve visually similar medical images from past patients (e.g. tissue from biopsies) to reference when making a medical decision with a new patient. However, no algorithm can perfectly capture an expert's ideal notion of similarity for every case: an image that is algorithmically determined to be similar may not be medically relevant to a doctor's specific diagnostic needs. In this paper, we identified the needs of pathologists when searching for similar images retrieved using a deep learning algorithm, and developed tools that empower users to cope with the search algorithm on-the-fly, communicating what types of similarity are most important at different moments in time. In two evaluations with pathologists, we found that these refinement tools increased the diagnostic utility of images found and increased user trust in the algorithm. The tools were preferred over a traditional interface, without a loss in diagnostic accuracy. We also observed that users adopted new strategies when using refinement tools, re-purposing them to test and understand the underlying algorithm and to disambiguate ML errors from their own errors. Taken together, these findings inform future human-ML collaborative systems for expert decision-making.
研究动机与目标
- 解決機器學習演算法在為病理學家檢索臨床相關相似醫學影像時的限制。
- 設計互動式工具,使病理學家能根據臨床判斷動態調整演算法相似性結果。
- 評估此類工具是否能提升診斷實用性、信任度以及對演算法決策的理解。
- 探討使用者在與不完美的機器學習系統互動時,如何調整其診斷策略。
提出的方法
- 開發互動式視覺化工具,使病理學家能在影像檢索過程中即時調整相似性標準。
- 整合反饋機制,使使用者能傳達哪些影像特徵對相似性具有臨床相關性。
- 設計工具協助使用者區分演算法錯誤與自身診斷錯誤。
- 透過兩項使用者研究,由實務病理學家參與,評估工具的易用性、診斷準確性與信任度。
- 使用質性與量化分析,評估使用者行為與決策策略的變化。
- 應用可解釋性技術(例如,注意力圖)協助使用者理解模型行為並進一步優化結果。
实验结果
研究问题
- RQ1病理學家在診斷決策過程中,如何有效應對深度學習模型產生的不完美相似性結果?
- RQ2互動式優化工具在多大程度上提升了檢索影像的診斷實用性?
- RQ3使用優化工具如何影響使用者對機器學習演算法的信任度?
- RQ4當與可適應、不完美的演算法互動時,病理學家會採用哪些新的診斷策略?
- RQ5優化工具是否能幫助使用者區分演算法錯誤與臨床錯誤?
主要发现
- 與標準演算法結果相比,優化工具顯著提升了檢索影像的診斷實用性。
- 使用者在使用互動式工具時,即使演算法產生不完美結果,仍報告對演算法有更高的信任度。
- 病理學家偏好優化介面,即使診斷準確性未受影響,也優於傳統靜態介面。
- 使用者採用新策略,例如測試模型行為,並使用工具區分演算法錯誤與臨床錯誤。
- 這些工具使使用者能重新利用系統進行模型探索,增強對演算法限制的理解。
- 系統支援更具協作性的「人機」工作流程,使使用者能根據專業知識主動塑造輸出結果。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。