[Paper Review] Explicability? Legibility? Predictability? Transparency? Privacy? Security? The Emerging Landscape of Interpretable Agent Behavior
The paper surveys and organizes a taxonomy of interpretability notions for agent behavior (explicability, legibility, predictability, transparency) and extends to privacy/security in both cooperative and adversarial settings, clarifying how observer models influence plan interpretation.
There has been significant interest of late in generating behavior of agents that is interpretable to the human (observer) in the loop. However, the work in this area has typically lacked coherence on the topic, with proposed solutions for "explicable", "legible", "predictable" and "transparent" planning with overlapping, and sometimes conflicting, semantics all aimed at some notion of understanding what intentions the observer will ascribe to an agent by observing its behavior. This is also true for the recent works on "security" and "privacy" of plans which are also trying to answer the same question, but from the opposite point of view -- i.e. when the agent is trying to hide instead of revealing its intentions. This paper attempts to provide a workable taxonomy of relevant concepts in this exciting and emerging field of inquiry.
Motivation & Objective
- Clarify coherent definitions and relationships among explicability, legibility, predictability, and transparency in agent behavior.
- Differentiate cooperative versus adversarial settings for interpretable planning.
- Explain how observer models and computation constraints affect plan interpretation.
- Highlight online versus offline interactions and their impact on interpretability measures.
Proposed method
- Present a general framework for modeling agent and observer (Pi^A, Pi^Theta) including planning problems, plans, computation models, and observation models.
- Define and distinguish explicability, predictability, legibility, and transparency within this framework.
- Discuss motion vs. task planning domains and the role of observer computational capabilities.
- Summarize related work and provide a table of concepts across cooperative and adversarial settings.
- Discuss open issues and potential extensions, including learning observer models and long-term interactions.
Experimental results
Research questions
- RQ1What are the precise definitions and relationships among explicability, legibility, predictability, and transparency in planning?
- RQ2How do observer models and computation constraints shape an agent’s ability to produce interpretable behavior?
- RQ3How do cooperative and adversarial settings alter the goals and methods for achieving interpretable or obfuscated plans?
- RQ4What are the key gaps and future directions in online versus offline interpretability and learning observer models?
Key findings
- A coherent taxonomy is proposed that links explicability, legibility, predictability, and transparency to observer models and goal/plan completion.
- Explicability and predictability are non-monotonic and can depend on plan prefixes or suffixes in online versus offline settings.
- Legibility or transparency concerns goal inference, requiring minimization of ambiguity across possible observer goals.
- The framework distinguishes motion and task planning, and discusses how observer computation power affects interpretability measures.
- There is discussion of adversarial settings (privacy, obfuscation, deception, security) and how they relate to the cooperative notions of interpretability.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.