[Paper Review] AIROGS: Artificial Intelligence for RObust Glaucoma Screening Challenge
The AIROGS challenge develops robust AI methods for glaucoma screening from color fundus photographs (CFPs), emphasizing ungradable input detection and real-world robustness across a large, diverse dataset. The top teams achieved performance comparable to expert clinicians and demonstrated strong generalization to external datasets.
The early detection of glaucoma is essential in preventing visual impairment. Artificial intelligence (AI) can be used to analyze color fundus photographs (CFPs) in a cost-effective manner, making glaucoma screening more accessible. While AI models for glaucoma screening from CFPs have shown promising results in laboratory settings, their performance decreases significantly in real-world scenarios due to the presence of out-of-distribution and low-quality images. To address this issue, we propose the Artificial Intelligence for Robust Glaucoma Screening (AIROGS) challenge. This challenge includes a large dataset of around 113,000 images from about 60,000 patients and 500 different screening centers, and encourages the development of algorithms that are robust to ungradable and unexpected input data. We evaluated solutions from 14 teams in this paper, and found that the best teams performed similarly to a set of 20 expert ophthalmologists and optometrists. The highest-scoring team achieved an area under the receiver operating characteristic curve of 0.99 (95% CI: 0.98-0.99) for detecting ungradable images on-the-fly. Additionally, many of the algorithms showed robust performance when tested on three other publicly available datasets. These results demonstrate the feasibility of robust AI-enabled glaucoma screening.
Motivation & Objective
- Assess the feasibility of robust AI-enabled glaucoma screening using CFPs in real-world, ungradable conditions.
- Create a large, diverse dataset and a challenge framework that promotes robustness to ungradable and unexpected inputs.
- Evaluate submitted algorithms on screening performance and input ungradability reliability, with external validation.
- Compare AI solutions against human experts and establish reproducibility via containerized submissions and public datasets.
Proposed method
- Provide a large, diverse training/testing dataset (112,732 CFPs from ~60,071 subjects across ~500 sites) with labels RG, NRG, or Ungradable.
- Require participants to submit containerized algorithms (Type 2 challenge) to ensure reproducibility and allow cloud-based evaluation on private test data.
- Evaluate solutions on two screening metrics (pAUC_S for RG at high specificity; SE@95SP_S) and two robustness metrics (kappa_U for ungradability agreement with humans; AUC_U for ungradability score correlation).
- Allow external validation by applying trained algorithms to three public datasets (REFUGE, GAMMA, DRIMDB) to assess generalization and robustness.
- Encourage methods that detect ungradable images on-the-fly without training on ungradable data.
Experimental results
Research questions
- RQ1Can AI models detect referable glaucoma from CFPs with high specificity and sensitivity in a real-world, unfiltered test set?
- RQ2Can AI systems reliably identify ungradable images and provide robust uncertainty measures in the presence of out-of-distribution data?
- RQ3Do AI solutions generalize well to external glaucoma datasets beyond the training domain?
- RQ4Is it feasible to achieve performance comparable to expert ophthalmologists using robust architectures and input-quality awareness?
Key findings
- The best teams achieved performance similar to a set of 20 expert ophthalmologists/optometrists on the glaucoma screening task.
- The top approach reached an AUC of 0.99 (95% CI: 0.98–0.99) for detecting ungradable images on-the-fly.
- Thirty teams participated across four challenge phases, with 14 teams contributing methods in the final paper.
- Algorithms demonstrated robust performance when evaluated on three external datasets (REFUGE, GAMMA, DRIMDB).
- The dataset is the largest public CFP glaucoma-label dataset to date, spanning 60k patients across 500 sites with diverse camera types.
- The challenge design (Type 2 submissions and unfiltered test set) enhances reproducibility and real-world relevance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.