[论文解读] Explicability? Legibility? Predictability? Transparency? Privacy? Security? The Emerging Landscape of Interpretable Agent Behavior
本文对代理人行为的可解释性概念(explicability、legibility、predictability、transparency)进行分类与整理,并扩展到合作与对抗场景中的隐私/安全性,阐明观测者模型如何影响对 plan 的解释。
There has been significant interest of late in generating behavior of agents that is interpretable to the human (observer) in the loop. However, the work in this area has typically lacked coherence on the topic, with proposed solutions for "explicable", "legible", "predictable" and "transparent" planning with overlapping, and sometimes conflicting, semantics all aimed at some notion of understanding what intentions the observer will ascribe to an agent by observing its behavior. This is also true for the recent works on "security" and "privacy" of plans which are also trying to answer the same question, but from the opposite point of view -- i.e. when the agent is trying to hide instead of revealing its intentions. This paper attempts to provide a workable taxonomy of relevant concepts in this exciting and emerging field of inquiry.
研究动机与目标
- 澄清代理人行为中 explicability、legibility、predictability 与 transparency 的清晰定义及相互关系。
- 区分用于可解释规划的合作与对抗设置。
- 解释观测者模型与计算约束如何影响对计划的解释。
- 强调在线与离线交互及其对可解释性度量的影响。
提出的方法
- 提出一个建模代理人与观测者( ѡPi^A, Pi^Theta)及其规划问题、计划、计算模型和观测模型的通用框架。
- 在该框架内定义并区分 explicability、predictability、legibility 和 transparency。
- 讨论运动规划与任务规划领域以及观测者的计算能力的作用。
- 总结相关工作并提供在合作与对抗设置下的概念表。
- 讨论未解决的问题与潜在扩展,包括学习观测者模型和长期交互。
实验结果
研究问题
- RQ1在规划中,explicability、legibility、predictability 和 transparency 的精确定义及相互关系是什么?
- RQ2观测者模型与计算约束如何影响代理人产生可解释行为的能力?
- RQ3合作与对抗设置如何改变实现可解释或混淆计划的目标与方法?
- RQ4在线与离线可解释性以及学习观测者模型的关键差距与未来方向是什么?
主要发现
- 提出了一个连贯的分类法,将 explicability、legibility、predictability 与 transparency 与观测者模型及目标/计划完成联系起来。
- explicability 与 predictability 是非单调的,并且在在线与离线设置下可能取决于计划前缀或后缀。
- Legibility 或 transparency 关注目标推断,需在可能的观测者目标之间最小化歧义。
- 该框架区分运动规划和任务规划,并讨论观测者的计算能力如何影响可解释性度量。
- 有关于对抗性设置(隐私、混淆、欺骗、安全性)及其与合作可解释性概念的关系的讨论。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。