[Paper Review] OdontoAI: A human-in-the-loop labeled data set and an online platform to boost research on dental panoramic radiographs
OdontoAI introduces a large-scale, human-in-the-loop (HITL)-labeled dataset of 4,000 dental panoramic radiographs with instance segmentation and tooth numbering, reducing labeling time by 51% (over 390 hours saved). The study also presents the OdontoAI online platform, the first benchmark platform for dental panoramic radiographs, enabling standardized evaluation of deep learning models in segmentation and tooth numbering tasks.
Deep learning has remarkably advanced in the last few years, supported by large labeled data sets. These data sets are precious yet scarce because of the time-consuming labeling procedures, discouraging researchers from producing them. This scarcity is especially true in dentistry, where deep learning applications are still in an embryonic stage. Motivated by this background, we address in this study the construction of a public data set of dental panoramic radiographs. Our objects of interest are the teeth, which are segmented and numbered, as they are the primary targets for dentists when screening a panoramic radiograph. We benefited from the human-in-the-loop (HITL) concept to expedite the labeling procedure, using predictions from deep neural networks as provisional labels, later verified by human annotators. All the gathering and labeling procedures of this novel data set is thoroughly analyzed. The results were consistent and behaved as expected: At each HITL iteration, the model predictions improved. Our results demonstrated a 51% labeling time reduction using HITL, saving us more than 390 continuous working hours. In a novel online platform, called OdontoAI, created to work as task central for this novel data set, we released 4,000 images, from which 2,000 have their labels publicly available for model fitting. The labels of the other 2,000 images are private and used for model evaluation considering instance and semantic segmentation and numbering. To the best of our knowledge, this is the largest-scale publicly available data set for panoramic radiographs, and the OdontoAI is the first platform of its kind in dentistry.
Motivation & Objective
- To address the scarcity of large-scale, consistently labeled dental panoramic radiograph datasets in deep learning research.
- To reduce the time and cost of labeling by implementing a human-in-the-loop (HITL) workflow using deep learning predictions as provisional labels.
- To create a public benchmark platform, OdontoAI, to standardize evaluation and comparison of deep learning models in dental image analysis.
- To improve model performance through iterative HITL refinement, where human experts verify and correct model predictions across multiple cycles.
- To support future research in precise tooth segmentation, numbering, and detection of dental structures like implants and prostheses.
Proposed method
- Employed a human-in-the-loop (HITL) pipeline where initial model predictions on unlabeled radiographs were verified and corrected by human dentists.
- Used deep neural networks (e.g., HTC, Mask R-CNN, Cascade R-CNN) to generate initial segmentation and numbering predictions for unlabeled images.
- Conducted iterative HITL cycles: trained models on verified data, generated new predictions, and repeated the verification process to improve label quality and model performance.
- Released 2,000 images with public labels for training and 2,000 with private labels for evaluation on the OdontoAI online platform.
- Implemented a comprehensive benchmarking system on the OdontoAI platform with metrics including mAP, exact match, micro-precision, micro-recall, and Hamming loss.
- Focused on fine-grained labeling of teeth, including instance segmentation and accurate numbering, to align with clinical diagnostic needs.
Experimental results
Research questions
- RQ1To what extent can the human-in-the-loop (HITL) approach reduce labeling time and effort in dental panoramic radiograph annotation?
- RQ2How does iterative HITL refinement improve the performance of deep learning models in tooth instance segmentation and numbering?
- RQ3What is the impact of using model-predicted labels as provisional annotations on the final label quality and consistency?
- RQ4How does the performance of state-of-the-art models (e.g., HTC, Mask R-CNN) compare on the new OdontoAI benchmark dataset?
- RQ5Can the OdontoAI platform serve as a reliable, standardized benchmark for evaluating and comparing deep learning models in dental panoramic radiograph analysis?
Key findings
- The HITL approach reduced labeling time by 51%, saving an estimated 390 continuous working hours in the annotation process.
- Model performance improved consistently across HITL iterations, with HTC achieving a 5.4 percentage point increase in segmentation mAP on the test set from iteration 1 to 4.
- The winning architecture, HTC, achieved an exact match score of 67.9% on the tooth numbering benchmark, with micro-precision and micro-recall above 98.5%.
- The OdontoAI dataset is the largest publicly available dataset for dental panoramic radiographs, with 4,000 images and detailed instance-level annotations.
- The platform supports fair model comparison and includes metrics such as Hamming loss (0.0143 for the top model), enabling comprehensive evaluation of multi-label prediction tasks.
- The study confirms that data collection via HITL is more efficient than model refinement when high-quality, general-purpose labels are required.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.