[Paper Review] A Survey of Mobile Computing for the Visually Impaired
This survey examines machine learning-based mobile applications for the visually impaired, evaluating their functionality in object detection, OCR, image captioning, and micro-navigation. It advocates for on-device processing via model compression and federated learning to enhance privacy, reduce latency, and improve accessibility, identifying key research directions in mobile perception, micro-navigation, and content summarization.
The number of visually impaired or blind (VIB) people in the world is estimated at several hundred million. Based on a series of interviews with the VIB and developers of assistive technology, this paper provides a survey of machine-learning based mobile applications and identifies the most relevant applications. We discuss the functionality of these apps, how they align with the needs and requirements of the VIB users, and how they can be improved with techniques such as federated learning and model compression. As a result of this study we identify promising future directions of research in mobile perception, micro-navigation, and content-summarization.
Motivation & Objective
- To analyze the current landscape of mobile assistive technologies for visually impaired or blind (VIB) users.
- To evaluate how well existing mobile applications meet the functional and usability needs of VIB individuals.
- To identify critical gaps in current systems, particularly in content summarization, micro-navigation, and contextual understanding.
- To advocate for on-device machine learning using model compression and federated learning to improve privacy, latency, and autonomy.
- To outline future research directions in mobile perception, micro-navigation, and intelligent content summarization for VIB users.
Proposed method
- Conducted interviews with VIB users and assistive technology developers to identify real-world needs and challenges.
- Surveyed existing mobile applications for object detection, OCR, image captioning, and scene understanding using on-device and server-based models.
- Evaluated model compression techniques including pruning, quantization, Huffman encoding, and knowledge distillation to reduce model size and energy use.
- Explored federated learning as a privacy-preserving method for collaborative model improvement without sharing raw user data.
- Analyzed the limitations of current datasets like MS COCO for VIB-specific needs and highlighted the value of community-driven datasets like VizWiz.
- Proposed integrating on-device processing for tasks like text extraction, color detection, and object recognition to reduce dependency on cloud services.
Experimental results
Research questions
- RQ1How do current mobile applications for the visually impaired perform in real-world tasks such as object detection, OCR, and image captioning?
- RQ2What are the key limitations of existing assistive technologies in terms of accuracy, latency, privacy, and user experience?
- RQ3How can model compression and knowledge distillation enable high-accuracy AI models to run efficiently on mobile devices?
- RQ4In what ways can federated learning improve model performance for VIB users while preserving data privacy?
- RQ5What are the most promising future research directions for mobile perception, micro-navigation, and content summarization in assistive technologies?
Key findings
- On-device models like BlindTool and SeeingAI demonstrate improved latency and privacy but are limited by model size and class generalization.
- Server-based captioning apps like SeeingAI suffer from slow response times and suboptimal interaction patterns, despite higher accuracy.
- Model compression techniques such as pruning, quantization, and knowledge distillation significantly reduce model size and energy consumption without major accuracy loss.
- Federated learning enables collaborative model improvement across users without exposing private images, making it ideal for VIB communities.
- Current datasets like MS COCO are insufficient for VIB needs, as they lack the detailed, context-rich descriptions required for effective navigation and task completion.
- There is a critical lack of integrated applications combining OCR and text summarization for physical documents, representing a major gap in current assistive technology.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.