[论文解读] OdontoAI: A human-in-the-loop labeled data set and an online platform to boost research on dental panoramic radiographs
OdontoAI 引入了一个大规模、人工介入(HITL)标注的数据集,包含 4,000 幅牙科全景X光片,具备实例分割与牙齿编号功能,将标注时间减少了 51%(节省超过 390 个工时)。该研究还推出了 OdontoAI 在线平台,这是首个针对牙科全景X光片的基准平台,支持深度学习模型在分割与牙齿编号任务中的标准化评估。
Deep learning has remarkably advanced in the last few years, supported by large labeled data sets. These data sets are precious yet scarce because of the time-consuming labeling procedures, discouraging researchers from producing them. This scarcity is especially true in dentistry, where deep learning applications are still in an embryonic stage. Motivated by this background, we address in this study the construction of a public data set of dental panoramic radiographs. Our objects of interest are the teeth, which are segmented and numbered, as they are the primary targets for dentists when screening a panoramic radiograph. We benefited from the human-in-the-loop (HITL) concept to expedite the labeling procedure, using predictions from deep neural networks as provisional labels, later verified by human annotators. All the gathering and labeling procedures of this novel data set is thoroughly analyzed. The results were consistent and behaved as expected: At each HITL iteration, the model predictions improved. Our results demonstrated a 51% labeling time reduction using HITL, saving us more than 390 continuous working hours. In a novel online platform, called OdontoAI, created to work as task central for this novel data set, we released 4,000 images, from which 2,000 have their labels publicly available for model fitting. The labels of the other 2,000 images are private and used for model evaluation considering instance and semantic segmentation and numbering. To the best of our knowledge, this is the largest-scale publicly available data set for panoramic radiographs, and the OdontoAI is the first platform of its kind in dentistry.
研究动机与目标
- 为解决深度学习研究中大规模、一致标注的牙科全景X光片数据集稀缺的问题。
- 通过实施利用深度学习预测作为临时标注的人工介入(HITL)工作流程,减少标注的时间与成本。
- 创建一个公开的基准平台 OdontoAI,以标准化牙科图像分析中深度学习模型的评估与比较。
- 通过迭代式HITL优化,提升模型性能,即人类专家在多个周期内验证并修正模型预测结果。
- 支持未来在精确牙齿分割、牙齿编号以及种植体和假牙等牙科结构检测方面的研究。
提出的方法
- 采用人工介入(HITL)流程,由牙科医生验证并修正未标注X光片的初始模型预测结果。
- 使用深度神经网络(如 HTC、Mask R-CNN、Cascade R-CNN)对未标注图像生成初始分割与编号预测。
- 开展迭代式HITL循环:基于已验证数据训练模型,生成新预测,并重复验证流程,以提升标注质量与模型性能。
- 在 OdontoAI 在线平台上发布 2,000 幅带有公开标签的图像用于训练,以及 2,000 幅带有私有标签的图像用于评估。
- 在 OdontoAI 平台上实施全面的基准测试系统,包含 mAP、精确匹配、微平均精确率、微平均召回率和汉明损失等指标。
- 专注于牙齿的细粒度标注,包括实例分割与准确编号,以契合临床诊断需求。
实验结果
研究问题
- RQ1人工介入(HITL)方法在牙科全景X光片标注中,能在多大程度上减少标注时间与工作量?
- RQ2迭代式HITL优化如何提升深度学习模型在牙齿实例分割与编号任务中的性能?
- RQ3使用模型预测标签作为临时标注,对最终标注质量与一致性有何影响?
- RQ4在新的 OdontoAI 基准数据集上,最先进模型(如 HTC、Mask R-CNN)的性能表现如何比较?
- RQ5OdontoAI 平台能否作为可靠、标准化的基准,用于评估与比较牙科全景X光片分析中的深度学习模型?
主要发现
- HITL 方法将标注时间减少了 51%,在标注过程中节省了约 390 个连续工作小时。
- 模型性能在 HITL 迭代过程中持续提升,HTC 在测试集上的分割 mAP 从第 1 轮到第 4 轮提升了 5.4 个百分点。
- 获胜架构 HTC 在牙齿编号基准测试中取得了 67.9% 的精确匹配得分,微平均精确率与微平均召回率均超过 98.5%。
- OdontoAI 数据集是目前公开可用的最大牙科全景X光片数据集,包含 4,000 幅图像及详细的实例级标注。
- 该平台支持公平的模型比较,并包含汉明损失(顶级模型为 0.0143)等指标,可全面评估多标签预测任务。
- 本研究证实,当需要高质量、通用性标注时,通过 HITL 收集数据比通过模型优化更高效。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。