[论文解读] Gaze Beyond Limits: Integrating Eye-Tracking and Augmented Reality for Next-Generation Spacesuit Interaction
本文提出了一种新颖的、即插即用的基于摄像头的方法,通过基于外观的注视估计,在日常手机交互中实现对视觉注意力的鲁棒感知与量化。通过整合多任务CNN、带有卡尔曼滤波的沙漏网络以及自适应阈值处理,该方法在多种移动场景下实现了眼接触检测的最先进性能,为真实世界人机交互研究提供了新的注意力度量指标。
Extravehicular activities (EVAs) are increasingly frequent in human spaceflight, particularly in spacecraft maintenance, scientific research, and planetary exploration. Spacesuits are essential for sustaining astronauts in the harsh environment of space, making their design a key factor in the success of EVA missions. The development of spacesuit technology has traditionally been driven by highly engineered solutions focused on life support, mission adaptability and operational efficiency. Modern spacesuits prioritize maintaining optimal internal temperature, humidity and pressure, as well as withstanding extreme temperature fluctuations and providing robust protection against micrometeoroid impacts and space debris. However, their bulkiness and rigidity impose significant physical strain on astronauts, reducing mobility and dexterity, particularly in tasks requiring fine motor control. The restricted field of view further complicates situational awareness, increasing the cognitive load during high-precision operations. While traditional spacesuits support basic EVA tasks, future space exploration shifting toward long-duration lunar and Martian surface missions demand more adaptive, intelligent, and astronaut-centric designs to overcome current constraints. To explore a next-generation spacesuit, this paper proposed an in-process eye-tracking embedded Augmented Reality (AR) Spacesuit System to enhance astronaut-environment interactions. By leveraging Segment-Anything Models (SAM) and Vision-Language Models (VLMs), we demonstrate a four-step approach to enable top-down gaze detection to minimize erroneous fixation data, gaze-based segmentation of objects of interest, real-time contextual assistance via AR overlays and hands-free operation within the spacesuit. This approach enhances real-time situational awareness and improves EVA task efficiency. We conclude with an exploration of the AR Helmet System’s potential in revolutionizing human-space interaction paradigms for future long-duration deep-space missions and discuss the further optimization of eye-tracking interactions using VLMs to predict astronaut intent and highlight relevant objects preemptively.
研究动机与目标
- 开发一种在自然、日常手机交互中无感感知视觉注意力的方法。
- 克服以往依赖专用眼动仪或触控事件等代理指标的方法的局限性。
- 实现在无需人工标注或外部硬件的情况下的准确、自动、原位注意力感知。
- 引入新的、可操作的注意力度量指标,用于研究移动人机交互中的注意力分配。
提出的方法
- 利用移动设备的前置摄像头进行基于外观的注视估计,消除了对专用眼动追踪硬件的需求。
- 采用多任务CNN以在动态移动环境中实现对部分可见面部的鲁棒人脸检测,这对实际应用至关重要。
- 将最先进的沙漏神经网络与卡尔曼滤波结合,以在高姿态变化条件下提升面部关键点检测与头部姿态估计的性能。
- 应用图像归一化,并在大规模GazeCapture数据集上训练注视估计器,以提升泛化能力。
- 采用自适应阈值处理以应对极端头部姿态下的不可靠注视估计,在必要时以头部姿态作为代理。
- 对视频数据进行离线处理,以提取眼接触事件并推导出更高阶的注意力度量指标,如注视次数和注意力持续时间。
实验结果
研究问题
- RQ1仅使用内置的前置摄像头,能否在自然、日常的手机交互中以高精度和鲁棒性检测眼接触?
- RQ2与现有方法相比,该方法在多种移动设备、用户和环境条件下表现如何?
- RQ3从连续的眼接触检测中可以推导出哪些新的注意力度量指标,以量化真实场景下的注意力分配?
- RQ4头部姿态变化在多大程度上影响注视估计的准确性?在移动场景中如何缓解这一问题?
主要发现
- 所提出的方法在两个公开可用的数据集上显著优于当前最先进的眼接触检测方法,表现出对设备、用户和环境变化的更强鲁棒性。
- 该方法可推导出新型注意力度量指标,如注视次数、注意力转移次数、平均注意力持续时间以及主要注意力焦点。
- 基于头部姿态的自适应阈值处理提升了极端姿态下的性能,减少了对可能不准确的注视估计的依赖。
- 该方法支持对自然状态下移动交互中注意力分配的离线分析,为人工标注或事件代理提供了完全自动且无代理的替代方案。
- 该方法为未来在移动设备上的实时部署铺平了道路,开启了中断预测、参与度测量和上下文感知界面等新应用。
- 该方法揭示了移动使用中注意力高度碎片化,平均注意力持续时间仅几秒钟,且在设备与环境之间频繁切换。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。