Skip to main content
QUICK REVIEW

[论文解读] ACROBAT -- a multi-stain breast cancer histological whole-slide-image data set from routine diagnostics for computational pathology

Philippe Weitz, Masi Valkonen|arXiv (Cornell University)|Nov 24, 2022
AI in cancer detection被引用 4
一句话总结

ACROBAT 是一个大规模、公开可用的全切片图像(WSI)数据集,包含来自1,153名原发性乳腺癌患者的4,212对匹配的苏木精-伊红(H&E)和免疫组织化学(IHC)染色全切片图像,数据源自常规诊断流程。该数据集支持计算病理学研究,涵盖图像配准、染色引导学习、虚拟染色及伪影检测,并通过13名标注者手动标注的37,000对特征点实现性能基准测试。

ABSTRACT

The analysis of FFPE tissue sections stained with haematoxylin and eosin (H&E) or immunohistochemistry (IHC) is an essential part of the pathologic assessment of surgically resected breast cancer specimens. IHC staining has been broadly adopted into diagnostic guidelines and routine workflows to manually assess status and scoring of several established biomarkers, including ER, PGR, HER2 and KI67. However, this is a task that can also be facilitated by computational pathology image analysis methods. The research in computational pathology has recently made numerous substantial advances, often based on publicly available whole slide image (WSI) data sets. However, the field is still considerably limited by the sparsity of public data sets. In particular, there are no large, high quality publicly available data sets with WSIs of matching IHC and H&E-stained tissue sections. Here, we publish the currently largest publicly available data set of WSIs of tissue sections from surgical resection specimens from female primary breast cancer patients with matched WSIs of corresponding H&E and IHC-stained tissue, consisting of 4,212 WSIs from 1,153 patients. The primary purpose of the data set was to facilitate the ACROBAT WSI registration challenge, aiming at accurately aligning H&E and IHC images. For research in the area of image registration, automatic quantitative feedback on registration algorithm performance remains available through the ACROBAT challenge website, based on more than 37,000 manually annotated landmark pairs from 13 annotators. Beyond registration, this data set has the potential to enable many different avenues of computational pathology research, including stain-guided learning, virtual staining, unsupervised pre-training, artefact detection and stain-independent models.

研究动机与目标

  • 为解决现有公开全切片图像数据集中,来自常规临床诊断流程的、具有匹配H&E与IHC染色的大规模高质量数据稀缺问题。
  • 支持图像配准算法的开发与基准测试,以实现H&E与IHC全切片图像之间的对齐。
  • 支持多样化的计算病理学研究,包括染色引导学习、虚拟染色及染色无关模型训练。
  • 通过37,000对人工标注的特征点提供标准化的配准性能基准。

提出的方法

  • 从多个临床机构收集女性原发性乳腺癌患者手术切除标本的全切片图像。
  • 为每位患者获取匹配的H&E与IHC染色全切片图像,确保诊断质量与临床相关性。
  • 由13名标注者人工标注37,000对特征点,以实现对图像配准算法的定量评估。
  • 通过arXiv和ACROBAT挑战赛网站发布数据集,确保可访问性与可复现性。
  • 设计数据集以支持配准之外的多种计算病理学任务,包括无监督预训练与伪影检测。
  • 实施标准化的数据整理与质量控制,以确保研究应用中的一致性与可靠性。

实验结果

研究问题

  • RQ1基于此数据集,深度学习模型能否实现H&E与IHC全切片图像之间准确且鲁棒的配准?
  • RQ2染色引导学习在多大程度上可提升模型在不同染色方案下的泛化能力?
  • RQ3当在该多染色、临床采集的全切片图像数据集上训练时,虚拟染色方法的效能如何?
  • RQ4在此数据集上进行无监督预训练,能否提升下游生物标志物评分任务的分类性能?
  • RQ5在真实临床诊断全切片图像上开发染色无关模型时,面临的主要挑战是什么?

主要发现

  • ACROBAT数据集包含来自1,153名患者的4,212张全切片图像,具有匹配的H&E与IHC染色,是目前乳腺癌病理学中规模最大、公开可用的多染色全切片图像数据集。
  • 数据集包含由13名标注者人工标注的37,000对特征点,可实现对图像配准算法的高精度基准测试。
  • 该数据集支持多样化的计算病理学应用,包括染色引导学习、虚拟染色与伪影检测。
  • 数据源自常规诊断工作流程,确保了高度的临床相关性与真实世界中的变异特性。
  • 数据集通过arXiv与ACROBAT挑战赛网站公开获取,性能指标可供算法评估使用。
  • 该数据集通过在不同染色模式下提供一致的解剖对应关系,支持染色无关模型的开发。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。