[Paper Review] Athena: Constructing Dialogues Dynamically with Discourse Constraints
Athena is a dynamic, topic-agnostic dialogue system that uses discourse constraints to orchestrate responses from multiple, modular response generators (RGs), enabling flexible, real-time conversation on open-domain topics. By avoiding static dialogue graphs and leveraging dynamic knowledge sources, Athena achieved higher user satisfaction in user-driven topics and reduced sensitivity to user personality and expectations through targeted system improvements.
This report describes Athena, a dialogue system for spoken conversation on popular topics and current events. We develop a flexible topic-agnostic approach to dialogue management that dynamically configures dialogue based on general principles of entity and topic coherence. Athena's dialogue manager uses a contract-based method where discourse constraints are dispatched to clusters of response generators. This allows Athena to procure responses from dynamic sources, such as knowledge graph traversals and feature-based on-the-fly response retrieval methods. After describing the dialogue system architecture, we perform an analysis of conversations that Athena participated in during the 2019 Alexa Prize Competition. We conclude with a report on several user studies we carried out to better understand how individual user characteristics affect system ratings.
Motivation & Objective
- To address the scalability and rigidity of traditional, hand-crafted dialogue flow graphs in open-domain conversational agents.
- To develop a flexible, topic-agnostic dialogue management architecture that supports dynamic, real-time response generation from diverse sources.
- To improve user satisfaction by reducing dependency on user personality and prior expectations through enhanced system coverage and responsiveness.
- To evaluate how user-driven vs. system-directed topics affect conversation quality and system ratings.
- To investigate the impact of user expectations and personality on perceived system quality in spoken dialogue systems.
Proposed method
- A dialogue manager dispatches discourse constraints to clusters of response generators (RGs), enabling dynamic, interleaved response selection during conversation.
- The system uses a contract-based architecture where RGs are independently triggered based on discourse-level rules, not pre-defined state transitions.
- Response generators include dynamic sources such as real-time knowledge graph traversals, feature-based retrieval, and modular conversation flows.
- A novel named entity resolution system integrates a large knowledge base with ensemble entity linking to improve grounding in topics.
- User studies were conducted with 32 and 54 participants to compare user-driven vs. system-directed topics and assess the impact of expectations and personality.
- System improvements focused on deepening coverage of core topics (e.g., dinosaurs, movies) to enhance engagement and reduce user bias.
Experimental results
Research questions
- RQ1How does user preference for topic control affect ratings of a dialogue system in open-domain conversations?
- RQ2To what extent do user personality traits and prior expectations influence perceived quality of a conversational agent?
- RQ3Can improving system coverage of core topics reduce sensitivity to user demographics and expectations?
- RQ4What is the impact of dynamic, constraint-based dialogue management on conversation quality and flexibility?
- RQ5How does interleaving responses from multiple, independent response generators enhance dialogue richness and coherence?
Key findings
- Users rated conversations significantly higher when they chose their own topics (p = 0.02), contrary to initial expectations that system-directed topics would score better.
- Extraverts (p = 0.019) and more conscientious users (p = 0.003) rated the system more highly overall, indicating personality influences perception.
- Users with higher initial expectations rated the system lower after use (p = 0.015), suggesting unrealistic expectations negatively impact satisfaction.
- After system improvements focused on core topics, overall ratings for system-directed conversations improved significantly (p = 0.04), while user-topic ratings remained stable (p = .99).
- Perceived user control over the system increased by 51% (p = 0.0001), and both personality and expectation effects disappeared, indicating greater system robustness.
- The system achieved a 17% improvement in overall ratings for system topics, demonstrating that enhanced coverage reduces user demographic bias.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.