Skip to main content
QUICK REVIEW

[Paper Review] When Eye-Tracking Meets Machine Learning: A Systematic Review on Applications in Medical Image Analysis

Sahar Moradizeyveh, Mehnaz Tabassum|arXiv (Cornell University)|Mar 12, 2024
Retinal Imaging and Analysis4 citations
TL;DR

This systematic review synthesizes the integration of eye-gaze tracking with machine learning (ML) and deep learning (DL) in medical image analysis, demonstrating how radiologists' visual attention patterns enhance model interpretability and diagnostic accuracy. By leveraging gaze data as a supervisory signal, the study highlights attention consistency and two-stream architectures that improve lesion detection and align AI decisions with human cognition.

ABSTRACT

Eye-gaze tracking research offers significant promise in enhancing various healthcare-related tasks, above all in medical image analysis and interpretation. Eye tracking, a technology that monitors and records the movement of the eyes, provides valuable insights into human visual attention patterns. This technology can transform how healthcare professionals and medical specialists engage with and analyze diagnostic images, offering a more insightful and efficient approach to medical diagnostics. Hence, extracting meaningful features and insights from medical images by leveraging eye-gaze data improves our understanding of how radiologists and other medical experts monitor, interpret, and understand images for diagnostic purposes. Eye-tracking data, with intricate human visual attention patterns embedded, provides a bridge to integrating artificial intelligence (AI) development and human cognition. This integration allows novel methods to incorporate domain knowledge into machine learning (ML) and deep learning (DL) approaches to enhance their alignment with human-like perception and decision-making. Moreover, extensive collections of eye-tracking data have also enabled novel ML/DL methods to analyze human visual patterns, paving the way to a better understanding of human vision, attention, and cognition. This systematic review investigates eye-gaze tracking applications and methodologies for enhancing ML/DL algorithms for medical image analysis in depth.

Motivation & Objective

  • To investigate how eye-gaze tracking data can improve the interpretability and performance of ML/DL models in medical image analysis.
  • To identify methodological gaps in integrating human visual attention with AI, particularly in 3D imaging and multimodal data fusion.
  • To examine the role of gaze data in training models that mimic human diagnostic reasoning and attention patterns.
  • To evaluate the impact of gaze-informed supervision on model robustness, especially in reducing false-negative errors in lesion detection.
  • To address challenges in data collection, annotation, and real-time processing of eye-tracking data in clinical settings.

Proposed method

  • Conducted a systematic review of peer-reviewed studies on eye-gaze tracking and ML/DL in medical imaging, focusing on data sources, methodologies, and evaluation metrics.
  • Classified studies based on their use of gaze data: either as a supervisory signal during training (e.g., attention consistency) or as a separate modality in two-stream architectures.
  • Analyzed key techniques such as attention consistency, where model attention is regularized to match human fixation maps using loss functions like cross-entropy or MSE.
  • Evaluated model performance using metrics including AUC, accuracy, F1-score, IoU, structural similarity, and peak signal-to-noise ratio.
  • Reviewed architectures including Vision Transformers (ViTs), Graph Neural Networks (GNNs), and Affine Transformer Networks for gaze-aware image processing.
  • Assessed the use of gaze data in both training/validation (with supervision) and inference (without gaze), identifying potential robustness trade-offs.

Experimental results

Research questions

  • RQ1How can eye-gaze tracking data improve the interpretability and accuracy of ML/DL models in medical image analysis?
  • RQ2What are the dominant architectural approaches for integrating gaze data with image data in medical AI systems?
  • RQ3To what extent does aligning model attention with human visual fixation patterns enhance diagnostic performance?
  • RQ4What are the key challenges in collecting, annotating, and utilizing eye-tracking data in clinical and research settings?
  • RQ5How do gaze-informed models compare to conventional approaches in detecting lesions across different imaging modalities?

Key findings

  • Eye-gaze data significantly improves model interpretability by aligning artificial attention with human visual search patterns, especially in lesion detection tasks.
  • Attention consistency architectures that use gaze data during training show improved alignment between model and human attention, though performance may drop during inference when gaze is not available.
  • Two-stream architectures, which process gaze and image data separately, offer better interpretability but are computationally expensive and less efficient.
  • Despite progress, most studies focus on 2D medical images, leaving a critical gap in the application of gaze-aware models to 3D imaging and volumetric analysis.
  • There is a notable lack of integration between gaze data and other clinical knowledge sources such as diagnostic criteria and radiology reports, limiting multimodal learning potential.
  • The use of gaze data as a supervisory signal enhances model robustness and reduces false-negative errors in nodule detection, particularly in chest X-rays and CT scans.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.