Skip to main content
QUICK REVIEW

[论文解读] Temporal-Difference Learning to Assist Human Decision Making during the Control of an Artificial Limb

Ann L. Edwards, Alexandra Kearney|arXiv (Cornell University)|Sep 18, 2013
Muscle activation and electromyography studies参考文献 12被引用 8
一句话总结

本文提出使用时序差分学习与通用价值函数(GVFs)来预测肌电控制机械臂用户切换控制模式的时间,使设备能够提前预判并主动建议最优控制转换。系统通过用户行为实时学习,借助预测性预测减少手动切换,并通过自然交互实现无奖励的直观用户修正,显著简化了辅助假肢中的人机控制。

ABSTRACT

In this work we explore the use of reinforcement learning (RL) to help with human decision making, combining state-of-the-art RL algorithms with an application to prosthetics. Managing human-machine interaction is a problem of considerable scope, and the simplification of human-robot interfaces is especially important in the domains of biomedical technology and rehabilitation medicine. For example, amputees who control artificial limbs are often required to quickly switch between a number of control actions or modes of operation in order to operate their devices. We suggest that by learning to anticipate (predict) a user's behaviour, artificial limbs could take on an active role in a human's control decisions so as to reduce the burden on their users. Recently, we showed that RL in the form of general value functions (GVFs) could be used to accurately detect a user's control intent prior to their explicit control choices. In the present work, we explore the use of temporal-difference learning and GVFs to predict when users will switch their control influence between the different motor functions of a robot arm. Experiments were performed using a multi-function robot arm that was controlled by muscle signals from a user's body (similar to conventional artificial limb control). Our approach was able to acquire and maintain forecasts about a user's switching decisions in real time. It also provides an intuitive and reward-free way for users to correct or reinforce the decisions made by the machine learning system. We expect that when a system is certain enough about its predictions, it can begin to take over switching decisions from the user to streamline control and potentially decrease the time and effort needed to complete tasks. This preliminary study therefore suggests a way to naturally integrate human- and machine-based decision making systems.

研究动机与目标

  • 通过最小化控制过程中的手动模式切换,减轻截肢者使用假肢时的认知与身体负担。
  • 开发一种实时自适应系统,在用户采取明确动作前预判其控制意图。
  • 在无需显式人类强化或奖励信号的情况下,将机器学习预测结果集成到人机交互中。
  • 探索预测建模如何支持辅助机器人中的自然、直观及半自主控制。

提出的方法

  • 系统采用时序差分学习训练通用价值函数(GVFs),基于实时肌电图(EMG)信号预测未来的用户控制决策。
  • GVFs 用于预测模式切换的时间以及用户下一步意图操作的关节,时间范围为 10 步。
  • 预测在持续交互过程中增量更新,实现对用户行为的实时适应。
  • 系统将用户发起的操作作为隐式反馈——选择某功能即确认预测,而选择其他选项则降低其置信度。
  • 该方法依赖于从 EMG 活动和任务情境中提取的状态表征,以建模用户意图与切换模式。
  • 无需显式奖励或示范;用户行为本身即作为模型优化的监督信号。

实验结果

研究问题

  • RQ1使用 GVF 的时序差分学习能否在多用途机械臂控制系统中准确预测用户模式切换的时间?
  • RQ2系统基于 EMG 信号与上下文信息,能否有效预判用户下一步意图操作的关节?
  • RQ3系统能否在无需显式强化或示范的情况下,提供自然且无奖励的反馈?
  • RQ4预测模型在多大程度上可减少人工肢体控制中的人工切换操作?
  • RQ5系统在不同会话和用户之间如何保持预测的准确性?

主要发现

  • 系统成功实现实时预判用户切换决策,预测值在实际用户发起切换前即开始上升。
  • 关节活动预测也提前上升,表明对用户意图的准确预测。
  • 在仅经过五次时序差分学习迭代后,系统在未见测试数据上仍保持了预测准确性。
  • 用户行为作为隐式反馈:选择某功能即确认预测,而选择其他选项则降低其置信度。
  • 该方法实现了自然、无奖励的用户校正或强化机制,无需显式训练信号。
  • 结果表明,预测建模可减少甚至消除对人工切换的需求,从而显著节省任务完成的时间与精力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。