Skip to main content
QUICK REVIEW

[论文解读] When Eye-Tracking Meets Machine Learning: A Systematic Review on Applications in Medical Image Analysis

Sahar Moradizeyveh, Mehnaz Tabassum|arXiv (Cornell University)|Mar 12, 2024
Retinal Imaging and Analysis被引用 4
一句话总结

本篇系统性综述综合分析了眼动追踪与机器学习(ML)及深度学习(DL)在医学影像分析中的融合应用,展示了放射科医生视觉注意力模式如何提升模型的可解释性与诊断准确性。通过将眼动数据作为监督信号,本研究强调了注意力一致性与双流架构在提升病灶检测能力、并使AI决策与人类认知保持一致方面的优势。

ABSTRACT

Eye-gaze tracking research offers significant promise in enhancing various healthcare-related tasks, above all in medical image analysis and interpretation. Eye tracking, a technology that monitors and records the movement of the eyes, provides valuable insights into human visual attention patterns. This technology can transform how healthcare professionals and medical specialists engage with and analyze diagnostic images, offering a more insightful and efficient approach to medical diagnostics. Hence, extracting meaningful features and insights from medical images by leveraging eye-gaze data improves our understanding of how radiologists and other medical experts monitor, interpret, and understand images for diagnostic purposes. Eye-tracking data, with intricate human visual attention patterns embedded, provides a bridge to integrating artificial intelligence (AI) development and human cognition. This integration allows novel methods to incorporate domain knowledge into machine learning (ML) and deep learning (DL) approaches to enhance their alignment with human-like perception and decision-making. Moreover, extensive collections of eye-tracking data have also enabled novel ML/DL methods to analyze human visual patterns, paving the way to a better understanding of human vision, attention, and cognition. This systematic review investigates eye-gaze tracking applications and methodologies for enhancing ML/DL algorithms for medical image analysis in depth.

研究动机与目标

  • 探究眼动追踪数据如何提升医学影像分析中机器学习/深度学习模型的可解释性与性能。
  • 识别在3D成像与多模态数据融合中,将人类视觉注意力与AI整合时存在的方法论空白。
  • 考察眼动数据在训练模拟人类诊断推理与注意力模式的模型中的作用。
  • 评估基于眼动信息的监督对模型鲁棒性的影响,特别是减少病灶检测中假阴性错误的效果。
  • 解决在临床环境中收集、标注及实时处理眼动数据所面临的挑战。

提出的方法

  • 对同行评审的关于眼动追踪与医学影像中机器学习/深度学习的研究进行了系统性综述,重点关注数据来源、方法论与评估指标。
  • 根据眼动数据的使用方式对研究进行分类:一种是在训练过程中作为监督信号(例如注意力一致性),另一种是作为双流架构中的独立模态。
  • 分析关键技术,如注意力一致性,即通过交叉熵或均方误差(MSE)等损失函数,对模型注意力进行正则化以匹配人类注视图。
  • 使用AUC、准确率、F1分数、交并比(IoU)、结构相似性及信噪比峰值(PSNR)等指标评估模型性能。
  • 回顾了包括视觉变换器(ViTs)、图神经网络(GNNs)及仿射变换器网络在内的架构,用于实现眼动感知的图像处理。
  • 评估眼动数据在训练/验证阶段(有监督)与推理阶段(无眼动)的应用,识别潜在的鲁棒性权衡。

实验结果

研究问题

  • RQ1眼动追踪数据在多大程度上能提升医学影像分析中机器学习/深度学习模型的可解释性与准确性?
  • RQ2在医学AI系统中,将眼动数据与图像数据融合的主流架构方法有哪些?
  • RQ3将模型注意力与人类视觉注视模式对齐,在多大程度上能提升诊断性能?
  • RQ4在临床与研究环境中,收集、标注与利用眼动数据面临哪些关键挑战?
  • RQ5基于眼动信息的模型在不同成像模态中检测病灶方面,与传统方法相比表现如何?

主要发现

  • 眼动数据显著提升了模型的可解释性,通过使人工注意力与人类视觉搜索模式对齐,尤其在病灶检测任务中表现突出。
  • 采用眼动数据进行训练的注意力一致性架构,可提升模型与人类注意力的一致性,但当推理阶段无法获取眼动数据时,性能可能下降。
  • 双流架构通过分别处理眼动与图像数据,提供了更好的可解释性,但计算成本较高,效率较低。
  • 尽管已有进展,但多数研究仍集中于2D医学图像,导致在3D成像与体积分析中应用眼动感知模型方面存在显著空白。
  • 眼动数据与其它临床知识源(如诊断标准与放射科报告)的整合明显不足,限制了多模态学习的潜力。
  • 将眼动数据用作监督信号可增强模型鲁棒性,并显著减少在肺部X光片与CT扫描中结节检测的假阴性错误。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。