Skip to main content
QUICK REVIEW

[Paper Review] Deep Learning--Based Scene Simplification for Bionic Vision

Nicole Han, Sudhanshu Srivastava|arXiv (Cornell University)|Jan 30, 2021
Neuroscience and Neural Engineering41 references4 citations
TL;DR

This paper proposes a deep learning-based scene simplification framework to enhance prosthetic vision for retinal implants by preprocessing natural scenes into simplified representations. Using a biologically realistic model of phosphene perception, it demonstrates that object segmentation significantly improves scene understanding in virtual patients compared to saliency or depth-based methods, with performance degrading as phosphene size and elongation increase.

ABSTRACT

Retinal degenerative diseases cause profound visual impairment in more than 10 million people worldwide, and retinal prostheses are being developed to restore vision to these individuals. Analogous to cochlear implants, these devices electrically stimulate surviving retinal cells to evoke visual percepts (phosphenes). However, the quality of current prosthetic vision is still rudimentary. Rather than aiming to restore "natural" vision, there is potential merit in borrowing state-of-the-art computer vision algorithms as image processing techniques to maximize the usefulness of prosthetic vision. Here we combine deep learning--based scene simplification strategies with a psychophysically validated computational model of the retina to generate realistic predictions of simulated prosthetic vision, and measure their ability to support scene understanding of sighted subjects (virtual patients) in a variety of outdoor scenarios. We show that object segmentation may better support scene understanding than models based on visual saliency and monocular depth estimation. In addition, we highlight the importance of basing theoretical predictions on biologically realistic models of phosphene shape. Overall, this work has the potential to drastically improve the utility of prosthetic vision for people blinded from retinal degenerative diseases.

Motivation & Objective

  • To improve the utility of bionic vision for individuals with retinal degenerative diseases by leveraging computer vision algorithms as preprocessing for retinal implants.
  • To evaluate how different deep learning-based scene simplification strategies—semantic segmentation, saliency, and monocular depth estimation—affect scene understanding in simulated prosthetic vision (SPV).
  • To use a psychophysically validated, neurobiologically inspired model of phosphene shape to generate realistic SPV predictions for more accurate performance evaluation.
  • To systematically assess perceptual performance across varying phosphene sizes and shapes, moving beyond the assumption of isolated, circular phosphenes.
  • To provide a foundation for future real-time, edge-compatible image processing pipelines tailored to individual retinal implant systems.

Proposed method

  • The study employs state-of-the-art deep learning models for semantic segmentation, visual saliency, and monocular depth estimation to preprocess natural outdoor scenes.
  • A psychophysically validated computational model of retinal prostheses (pulse2percept) is used to simulate phosphene formation, incorporating realistic spatial spread and shape parameters (radius ρ and elongation λ).
  • The processed images are converted into simulated prosthetic vision (SPV) outputs by mapping segmented or salient regions to phosphene patterns based on electrode grid configurations.
  • A user study is conducted with sighted subjects (virtual patients) performing a binary detection task (identifying people and cars) in SPV-rendered scenes.
  • Performance is quantified using the sensitivity index d′, with statistical analysis comparing the impact of different simplification strategies and phosphene parameters.
  • The study evaluates electrode grid sizes (8×8, 16×16, 32×32) to assess the effect of spatial resolution on perceptual performance.

Experimental results

Research questions

  • RQ1Does deep learning-based scene simplification improve scene understanding in simulated prosthetic vision compared to baseline methods?
  • RQ2How do different scene simplification strategies—semantic segmentation, saliency, and depth estimation—compare in supporting object detection tasks?
  • RQ3To what extent does realistic phosphene shape (size and elongation) affect perceptual performance in SPV?
  • RQ4Does increasing the number of electrodes in the implant grid lead to measurable improvements in scene understanding?
  • RQ5Can biologically realistic phosphene modeling improve the validity of SPV predictions compared to idealized, circular phosphenes?

Key findings

  • Object segmentation significantly outperformed saliency and depth-based models in supporting scene understanding, with the highest d′ values across all conditions.
  • Performance degraded with increasing phosphene size (ρ) and elongation (λ), with the worst performance observed at ρ = 500 μm and λ = 2.0.
  • Even with large, elongated phosphenes, subjects performed above chance level (d′ > 0), indicating that all simplification strategies enabled detectable perception.
  • Increasing the electrode grid from 8×8 to 16×16 improved performance, but further increasing to 32×32 yielded no significant improvement.
  • The study confirms that realistic phosphene shape modeling is essential, as assuming isolated, circular phosphenes may lead to overly optimistic performance predictions.
  • The results suggest that semantic segmentation is the most effective preprocessing strategy for enhancing the utility of bionic vision in outdoor navigation tasks.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.