[Paper Review] Making Videos Accessible for Blind and Low Vision Users Using a Multimodal Agent Video Player
The paper presents a Multimodal Agent Video Player (MAVP) that uses a multimodal LLM and conversational architecture to give BLV users an interactive, autonomous video-watching experience, validated through user studies.
Video content remains largely inaccessible to blind and low-vision (BLV) users. To address this, we introduce a prototype that leverages a multimodal agent - powered by a novel conversational architecture using a multimodal large language model (MLLM) - to provide BLV users with an interactive, accessible video experience. This Multimodal Agent Video Player (MAVP) demonstrates that an interactive accessibility mode can be added to a video through multilayered prompt orchestration. We describe a user-centered design process involving 18 sessions with BLV users that showed that BLV users do not just want accessibility features, but desire independence and personal agency over the viewing experience. We conducted a qualitative study with an additional 8 BLV participants; in this, we saw that the MAVP's conversational dialogue offers BLV users a sense of personal agency, fostering collaboration and trust. Even in the case of hallucinations, it is meta-conversational dialogues about AI's limitations that can repair trust.
Motivation & Objective
- Demonstrate that an interactive accessibility mode can be added to video content through multimodal agent technology.
- Investigate how a multimodal large language model (MLLM) and conversational architecture support BLV users in controlling and understanding video content.
- Assess user perceived independence, agency, and trust in the MAVP system through qualitative studies with BLV participants.
- Explore how meta-conversational dialogues about AI limitations can repair trust in the presence of AI hallucinations.
Proposed method
- Develop a prototype MAVP that integrates a multimodal LLM for conversational interaction during video viewing.
- Employ multilayered prompt orchestration to create an interactive accessibility mode within the video player.
- Conduct a user-centered design process with 18 BLV sessions to gather qualitative insights on independence and agency.
- Conduct an additional qualitative study with 8 BLV participants to assess perceived agency, collaboration, and trust.
- Analyze how conversations about AI limitations affect trust and user satisfaction.
Experimental results
Research questions
- RQ1Can a multimodal agent video player provide BLV users with a sense of independence and personal agency over the viewing experience?
- RQ2Does the MAVP conversational dialogue foster collaboration and trust between BLV users and the AI assistant?
- RQ3How do meta-conversational discussions about AI limitations impact user trust during interactive video access?
- RQ4What design insights emerge from user sessions to improve accessibility and autonomy in video playback?
Key findings
- BLV users desire independence and personal agency, not merely accessibility features.
- MAVP’s conversational dialogue fosters a sense of collaboration and trust with BLV users.
- Meta-conversational discussions about AI limitations can help repair trust even when the system hallucinates.
- A user-centered design process with BLV participants yields actionable insights for interactive accessibility in video players.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.