Skip to main content
QUICK REVIEW

[논문 리뷰] Automatic Gaze Analysis: A Survey of Deep Learning based Approaches

Shreya Ghosh, Abhinav Dhall|arXiv (Cornell University)|2021. 08. 12.
Gaze Tracking and Assistive Technology참고 문헌 199인용 수 16
한 줄 요약

이 종합 검토는 자율적, 자기학습 및 약한 지도 학습 방법을 중점으로 하여 딥 러닝 기반의 자동 시선 분석 접근법을 포괄적으로 검토한다. 비구속 환경에서의 핵심 과제를 규명하고 AR/VR, HCI 및 컴퓨터 비전 응용 분야에서 강건하고 실시간 시선 추정을 위한 향후 방향을 제안한다.

ABSTRACT

Eye gaze analysis is an important research problem in the field of Computer Vision and Human-Computer Interaction. Even with notable progress in the last 10 years, automatic gaze analysis still remains challenging due to the uniqueness of eye appearance, eye-head interplay, occlusion, image quality, and illumination conditions. There are several open questions, including what are the important cues to interpret gaze direction in an unconstrained environment without prior knowledge and how to encode them in real-time. We review the progress across a range of gaze analysis tasks and applications to elucidate these fundamental questions, identify effective methods in gaze analysis, and provide possible future directions. We analyze recent gaze estimation and segmentation methods, especially in the unsupervised and weakly supervised domain, based on their advantages and reported evaluation metrics. Our analysis shows that the development of a robust and generic gaze analysis method still needs to address real-world challenges such as unconstrained setup and learning with less supervision. We conclude by discussing future research directions for designing a real-world gaze analysis system that can propagate to other domains including Computer Vision, Augmented Reality (AR), Virtual Reality (VR), and Human Computer Interaction (HCI). Project Page: https://github.com/i-am-shreya/EyeGazeSurvey}{https://github.com/i-am-shreya/EyeGazeSurvey

연구 동기 및 목표

  • 최근 딥 러닝 기반의 시선 추정 및 세분화 기술의 발전, 특히 약한 지도 및 비지도 설정에서의 발전을 분석하기 위해.
  • 제한된 지도 정보 하에 실제 세계의 비구속 조건에서 효과적인 시선 분석 기법을 규명하기 위해.
  • 시선 분석에 특화된 도메인별 메트릭 및 벤치마크 프로토콜을 사용하여 현재의 방법들을 평가하기 위해.
  • AR/VR 및 인간-컴퓨터 상호작용에서 강건하고 저지연 시선 추정을 위한 향후 연구 방향을 탐색하기 위해.
  • 기존 모델 기반 접근법과 데이터 기반 딥 러닝 간의 격차를 메우기 위해 하이브리드 학습 프레임워크를 제안하기 위해.

제안 방법

  • 딥 러닝을 활용한 시선 추정, 세분화 및 추적 분야의 최근 100편 이상의 논문에 대한 체계적 검토.
  • 지도 수준에 따라 방법을 분류: 완전히 지도 학습, 약한 지도 학습, 자기학습, 비지도 학습.
  • 핵심 구성 요소 분석: 눈 검출(등록), 특징 표현(예: CNN, ViT), 시선 예측(회귀 또는 분류).
  • 표준 벤치마크(예: CAVE, MPII, ETH-XGaze, TabletGaze)를 사용하여 평가하며, 평균 절대 오차(MAE) 및 각도 오차와 같은 메트릭을 활용.
  • RGB/IR 카메라, 노트북/웹캠, 전용 눈 추적 장치(예: 비디오 옥류그래피) 등의 데이터 수집 장치 종합 분석.
  • 컴퓨터 비전, AR/VR 및 HCI 분야의 통찰을 통합하여 기하학적 눈 모델과 딥 러닝 외관 특징을 융합한 하이브리드 모델 제안.
Figure 1: A brief chronology of seminal gaze analysis works. The very first gaze pattern modelling dates back to the work of Javal et al. in 1879 [ 4 ] . One of the first deep learning driven appearance based gaze estimation models was proposed in $2015$ [ 24 ] .
Figure 1: A brief chronology of seminal gaze analysis works. The very first gaze pattern modelling dates back to the work of Javal et al. in 1879 [ 4 ] . One of the first deep learning driven appearance based gaze estimation models was proposed in $2015$ [ 24 ] .

실험 결과

연구 질문

  • RQ1비구속 환경에서 시선 추정에 가장 효과적인 딥 러닝 아키텍처와 학습 파라다임은 무엇인가?
  • RQ2정확도와 일반화 능력 측면에서 비지도 및 약한 지도 학습 방법은 완전히 지도 학습 방법과 어떻게 비교되는가?
  • RQ3실제 세계의 시선 추정에서의 핵심 과제는 무엇이며, 현재의 방법들은 이를 어떻게 해결하는가?
  • RQ4단순한 시선 방향을 넘어서, 시선 추론을 통해 인지적 및 정서적 상태를 어떻게 추론할 수 있는가?
  • RQ5다중 모odal 또는 크로스 모달 입력(예: 음성, 머리 자세)이 시선 추정의 강건성 향상에 어떤 역할을 할 수 있는가?

주요 결과

  • 비지도 및 자기학습 방법은 비용이 많이 들고 오류가 발생하기 쉬운 인간 레이블 기반 시선 레이블에 대한 의존도를 줄이는 데 잠재력이 있다.
  • 현재 최상의 모델들은 제한된 조건에서 ETH-XGaze와 같은 벤치마크 데이터셋에서 평균 절대 오차가 1.5도 이하로 달성하고 있다.
  • 기하학적 눈 모델과 딥 러닝 외관 특징을 융합한 하이브리드 모델은 다양한 머리 자세와 조명 조건에서도 일반화 능력을 향상시킨다.
  • 향후 시선 궤적 예측은 AR/VR 응용 분야에서 저지연 푸코카라이징 렌더링을 가능하게 하는 핵심 기술로 부상하고 있다.
  • 음성 또는 머리 운동 신호와 같은 다중 모달 접근법은 가시성이 낮거나 가림이 발생하는 상황에서 시선 추정 성능을 향상시킬 수 있다.
  • 진전이 있었음에도 불구하고, 특히 극단적인 머리 자세와 가림이 발생하는 비구속 실세계 환경에서의 강건성은 여전히 주요 열린 과제로 남아 있다.
Figure 2: Top Left: Overview of the human visual system, eye modelling and eye movement. For computer vision based automated gaze analysis, we consider an image containing eyes (left) as input. Thus, such methods analyze the visible eye regions (middle) and predict the 2-D/3-D gaze vector as output.
Figure 2: Top Left: Overview of the human visual system, eye modelling and eye movement. For computer vision based automated gaze analysis, we consider an image containing eyes (left) as input. Thus, such methods analyze the visible eye regions (middle) and predict the 2-D/3-D gaze vector as output.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.