[论文解读] Teaching Social Behavior through Human Reinforcement for Ad hoc Teamwork -The STAR Framework
STAR框架通过为有效性与社会可接受性分别设立专用反馈通道,利用并行强化学习,教导临时组建的智能体团队遵守人类社会规范。结果表明,采用双重反馈时,智能体能够平衡高性能与社会合规行为,优于混合或单一反馈方法。
As AI technology continues to develop, more and more agents will become capable of long term autonomy alongside people. Thus, a recent line of research has studied the problem of teaching autonomous agents the concept of ethics and human social norms. Most existing work considers the case of an individual agent attempting to learn a predefined set of rules. In reality, however, social norms are not always pre-defined and are very difficult to represent algorithmically. Moreover, the basic idea behind the social norms concept is ensuring that one's actions do not negatively influence others' utilities, which is inherently a multiagent concept. Thus, here we investigate a way to teach agents, as a team, how to act according to human social norms. In this research, we introduce the STAR framework used to teach an ad hoc team of agents to act in accordance with human social norms. Using a hybrid team (agents and people), when taking an action considered to be socially unacceptable, the agents receive negative feedback from the human teammate(s) who has(have) an awareness of the team's norms. We view STAR as an important step towards teaching agents to act more consistently with respect to human morality.
研究动机与目标
- 解决在多智能体、人机协同环境中教导智能体社会规范的空白。
- 克服现有方法在单一智能体、预设规范或目标与约束形式混杂方面的局限。
- 通过团队协作过程中的实时人类反馈,使智能体学习到具有文化与时间相关性的社会规范。
- 开发一种支持长期自主性的框架,同时确保混合人机团队中的社会负责任行为。
- 验证为有效性与社会规范分别设立并行反馈通道,相较于混合或单通道反馈,能带来更优的学习效果。
提出的方法
- 提出STAR(通过强化学习进行社会性训练智能体)框架,采用两条并行反馈通道:一条用于任务有效性,一条用于社会可接受性。
- 由人类队友在智能体执行被认为社会不可接受的行为时,提供实时、专用的反馈,无论其有效性如何。
- 使用强化学习训练智能体,同时优化任务表现与社会规范的遵守。
- 设计混合团队环境,包含人类与智能体队友,以模拟现实世界中的临时团队协作场景。
- 采用双通道反馈机制,将社会规范与任务有效性解耦,避免学习过程中的相互干扰。
- 评估多种反馈设计:仅有效性、仅社会性、混合反馈,以及所提出的并行STAR框架。
实验结果
研究问题
- RQ1智能体能否在无预先指定社会规范的情况下,通过人类反馈在临时团队中学习到社会可接受的行为?
- RQ2与混合反馈相比,将有效性与社会规范反馈分离,对学习表现有何影响?
- RQ3STAR中的并行反馈设计是否能使智能体在保持高性能的同时,最大限度减少社会不可接受行为?
- RQ4该框架能否使智能体在不同文化或时间背景下泛化社会规范?
- RQ5在多智能体系统中,教学有效性和社会合规性的最优反馈结构是什么?
主要发现
- 并行STAR反馈设计在有效性与社会合规性两方面均最接近理论最优上界。
- 仅接收社会性反馈的智能体表现出最高比例的可接受行为,但游戏清除行数最少,表明有效性最差。
- 仅有效性反馈导致清除行数最多,但社会不可接受行为也最多。
- 混合反馈设计未能有效教授社会规范,尽管有效性表现高,但社会合规性表现曲线始终偏低。
- STAR框架使智能体能够并行学习两个目标,在高性能与社会可接受性之间实现平衡。
- 经过三轮游戏,STAR框架的性能曲线紧密跟随上界,证明了有效性与社会规范的联合学习效果显著。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。