[论文解读] Multimodal Fusion of EMG and Vision for Human Grasp Intent Inference in Prosthetic Hand Control
本文提出了一种贝叶斯证据融合框架,结合肌电信号(EMG)与视觉数据(包括视线追踪和RGB视频),以提升假肢手控制中的抓握意图推断。通过将基于神经网络的EMG分类器与视觉分类器结合,并引入时间相位分割,该方法在关键的抓握前伸阶段实现了95.3%的平均分类准确率,相较于单一模态分别提升了13.66%(EMG)和14.8%(视觉)。
Objective: For transradial amputees, robotic prosthetic hands promise to regain the capability to perform daily living activities. Current control methods based on physiological signals such as electromyography (EMG) are prone to yielding poor inference outcomes due to motion artifacts, muscle fatigue, and many more. Vision sensors are a major source of information about the environment state and can play a vital role in inferring feasible and intended gestures. However, visual evidence is also susceptible to its own artifacts, most often due to object occlusion, lighting changes, etc. Multimodal evidence fusion using physiological and vision sensor measurements is a natural approach due to the complementary strengths of these modalities. Methods: In this paper, we present a Bayesian evidence fusion framework for grasp intent inference using eye-view video, eye-gaze, and EMG from the forearm processed by neural network models. We analyze individual and fused performance as a function of time as the hand approaches the object to grasp it. For this purpose, we have also developed novel data processing and augmentation techniques to train neural network components. Results: Our results indicate that, on average, fusion improves the instantaneous upcoming grasp type classification accuracy while in the reaching phase by 13.66% and 14.8%, relative to EMG (81.64% non-fused) and visual evidence (80.5% non-fused) individually, resulting in an overall fusion accuracy of 95.3%. Conclusion: Our experimental data analyses demonstrate that EMG and visual evidence show complementary strengths, and as a consequence, fusion of multimodal evidence can outperform each individual evidence modality at any given time.
研究动机与目标
- 解决单模态控制在假肢手中的局限性,如EMG中的运动伪影和信号漂移,以及视觉中的遮挡或光照问题。
- 提升在抓握前伸阶段的抓握意图推断鲁棒性与准确性,该阶段决策时机对机器人执行至关重要。
- 开发一种融合框架,充分利用EMG与视觉数据的互补优势,以提升分类性能。
- 引入新型数据处理与增强技术,包括复制粘贴增强与时间相位分割,以提升模型泛化能力。
- 证明多模态融合在所有阶段均持续优于单一模态,尤其在抓握前伸阶段表现更优。
提出的方法
- 采用贝叶斯证据融合框架,利用最大似然估计融合EMG与视觉证据。
- 使用卷积神经网络(CNN)基于眼视图RGB视频与视线数据分类抓握类型,并通过复制粘贴增强实现背景泛化。
- 为EMG信号分类实现独立的神经网络,包含相位检测(静息、前伸、抓握、返回),以保留时间动态特性。
- 将每个抓握序列划分为不同的时间相位,避免不同阶段肌肉活动模式之间的混淆。
- 对融合决策应用时间平滑处理,通过利用历史决策与系统约束,将准确率提升至96.8%。
- 在下肢截肢者执行日常物体交互任务时同步采集的EMG与视觉数据集上训练并评估模型。
实验结果
研究问题
- RQ1与单一模态相比,EMG与视觉的多模态融合在抓握意图分类准确率方面有何提升?
- RQ2时间相位分割对基于EMG的抓握分类性能有何影响?
- RQ3在抓握-前伸周期的哪个阶段,融合能带来最大的性能增益?
- RQ4视觉与EMG模态能否相互补偿对方的局限性,如遮挡或静息期EMG信号微弱?
- RQ5决策平滑处理在多大程度上提升了融合抓握预测的鲁棒性与准确性?
主要发现
- 与仅使用EMG相比,EMG与视觉证据融合在前伸阶段将抓握分类准确率提升了13.66%。
- 与仅使用视觉相比,融合在前伸阶段将准确率提升了14.8%,表明性能有显著提升。
- 整体融合准确率达到95.3%,其中前伸阶段准确率最高(95.3%),该阶段对控制决策最为关键。
- 在静息阶段,视觉分类器表现优于EMG(90.69% vs. 16.86%),凸显了各模态在不同阶段的互补优势。
- 即使两种模态均提供强证据,融合仍进一步提升了准确率,表明其在性能上超越单一模态,具备额外鲁棒性。
- 对融合输出应用决策平滑处理后,平均准确率提升至96.8%,表明时间一致性可增强实际应用中的可用性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。