[论文解读] Robust Real-Time Multi-View Eye Tracking
本文提出了一种实时多相机眼动追踪框架,通过同时利用多个眼部外观来增强在真实环境条件下的鲁棒性。通过一种基于可靠性的自适应融合机制,结合头部姿态和预测的注视行为,该系统在30 fps下实现了约1°的精度,与单视角方法相比,显著提升了对头部运动、眼镜和光照变化的鲁棒性。
Despite significant advances in improving the gaze tracking accuracy under controlled conditions, the tracking robustness under real-world conditions, such as large head pose and movements, use of eyeglasses, illumination and eye type variations, remains a major challenge in eye tracking. In this paper, we revisit this challenge and introduce a real-time multi-camera eye tracking framework to improve the tracking robustness. First, differently from previous work, we design a multi-view tracking setup that allows for acquiring multiple eye appearances simultaneously. Leveraging multi-view appearances enables to more reliably detect gaze features under challenging conditions, particularly when they are obstructed in conventional single-view appearance due to large head movements or eyewear effects. The features extracted on various appearances are then used for estimating multiple gaze outputs. Second, we propose to combine estimated gaze outputs through an adaptive fusion mechanism to compute user's overall point of regard. The proposed mechanism firstly determines the estimation reliability of each gaze output according to user's momentary head pose and predicted gazing behavior, and then performs a reliability-based weighted fusion. We demonstrate the efficacy of our framework with extensive simulations and user experiments on a collected dataset featuring 20 subjects. Our results show that in comparison with state-of-the-art eye trackers, the proposed framework provides not only a significant enhancement in accuracy but also a notable robustness. Our prototype system runs at 30 frames-per-second (fps) and achieves 1 degree accuracy under challenging experimental scenarios, which makes it suitable for applications demanding high accuracy and robustness.
研究动机与目标
- 解决现有眼动追踪系统在真实环境条件下(如大幅头部运动、佩戴眼镜和光照变化)鲁棒性不足的问题。
- 克服单视角眼动追踪的局限性,即在挑战性条件下特征可能被遮挡或失真。
- 开发一种实时多相机框架,实现在多个眼部外观下可靠地检测特征。
- 设计一种自适应融合机制,根据从头部姿态和预测注视行为中获得的可靠性来加权注视估计。
- 在无需下巴托或复杂校准的情况下,证明在非约束环境中具有高精度和鲁棒性。
提出的方法
- 部署多视角相机系统,同时捕捉多个眼部外观,以在遮挡或失真情况下提高特征可见性。
- 从每个相机视角的眼部图像中独立提取与注视相关的特征(例如,瞳孔中心、反光点)。
- 在每个视角上使用注视估计方法(例如,基于交叉比或回归的方法)独立估计多个注视输出。
- 实施一种自适应融合机制,基于瞬时头部姿态和预测的注视行为,为每个注视估计计算可靠性评分。
- 使用可靠性评分对注视输出进行加权融合,生成单一、鲁棒的注视点估计。
- 优化系统以实现实时性能,在原型实现中以低分辨率(90×50)眼部图像实现30 fps。
实验结果
研究问题
- RQ1与单视角系统相比,多视角眼动追踪是否能提升在大幅头部运动下的注视估计鲁棒性?
- RQ2当因眼镜或头部运动导致遮挡时,使用多个眼部外观在多大程度上提高了特征检测的可靠性?
- RQ3一种考虑头部姿态和注视预测的自适应融合机制,在多大程度上提升了整体精度和可用性?
- RQ4低分辨率、未经校准的多相机设置是否能在无需下巴托或复杂校准的情况下实现高精度(≤1°)?
- RQ5所提出的框架在不同光照条件和多种眼型下的表现如何?
主要发现
- 与单相机设置相比,所提出的多视角框架在挑战性条件下将估计精度提高了约20%(0.2–0.6°),可用性提高了10–20%。
- 在真实场景中,系统实现了约1°的平均估计误差,且估计可用性接近100%,包括大幅头部运动和佩戴眼镜的情况。
- 在正常条件下,自适应融合机制通过根据可靠性动态加权注视估计,将精度提高了30%。
- 系统以30帧每秒运行,适用于虚拟现实/增强现实和人机交互等实时应用。
- 该框架对光照变化和个体间眼型差异具有高度容忍性,性能无显著下降。
- 原型系统无需下巴托,且使用低分辨率(90×50)眼部图像,证明即使在硬件要求极低的情况下也能实现高精度。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。