Skip to main content
QUICK REVIEW

[论文解读] Extracting Thyroid Nodules Characteristics from Ultrasound Reports Using Transformer-based Natural Language Processing Methods

Aman Pathak, Zehao Yu|PubMed|Mar 31, 2023
Topic Modeling参考文献 22被引用 7
一句话总结

本研究提出了一种基于Transformer的NLP框架,用于从超声报告中提取16种临床相关的甲状腺结节特征,采用精心筛选的语料库和微调模型(包括GatorTron)。GatorTron在严格标准下的F1得分达到0.8851,在宽松标准下达到0.9495,优于BERT、RoBERTa、LongFormer和DeBERTa,标志着首次系统性地将大规模NLP应用于临床文本中对甲状腺结节的全面特征表征。

ABSTRACT

The ultrasound characteristics of thyroid nodules guide the evaluation of thyroid cancer in patients with thyroid nodules. However, the characteristics of thyroid nodules are often documented in clinical narratives such as ultrasound reports. Previous studies have examined natural language processing (NLP) methods in extracting a limited number of characteristics (<9) using rule-based NLP systems. In this study, a multidisciplinary team of NLP experts and thyroid specialists, identified thyroid nodule characteristics that are important for clinical care, composed annotation guidelines, developed a corpus, and compared 5 state-of-the-art transformer-based NLP methods, including BERT, RoBERTa, LongFormer, DeBERTa, and GatorTron, for extraction of thyroid nodule characteristics from ultrasound reports. Our GatorTron model, a transformer-based large language model trained using over 90 billion words of text, achieved the best strict and lenient F1-score of 0.8851 and 0.9495 for the extraction of a total number of 16 thyroid nodule characteristics, and 0.9321 for linking characteristics to nodules, outperforming other clinical transformer models. To the best of our knowledge, this is the first study to systematically categorize and apply transformer-based NLP models to extract a large number of clinical relevant thyroid nodule characteristics from ultrasound reports. This study lays ground for assessing the documentation quality of thyroid ultrasound reports and examining outcomes of patients with thyroid nodules using electronic health records.

研究动机与目标

  • 从超声报告中识别并系统分类16种与临床决策相关的甲状腺结节特征。
  • 通过NLP专家与甲状腺专科医生的合作,制定标准化的标注指南。
  • 构建高质量、多学科参与的标注语料库,用于训练和评估NLP模型在甲状腺结节报告提取任务中的表现。
  • 比较最先进基于Transformer的模型在从非结构化临床叙述中提取大量甲状腺结节特征方面的性能。
  • 为未来基于电子健康记录的甲状腺结节管理结果研究和文档质量评估提供支持。

提出的方法

  • 由NLP专家与甲状腺专科医生组成的多学科团队,定义了16项与临床决策相关的甲状腺结节关键特征。
  • 依据标准化指南对精选的超声报告语料库进行标注,以确保一致性和临床相关性。
  • 在标注语料库上对五种最先进基于Transformer的模型(BERT、RoBERTa、LongFormer、DeBERTa和GatorTron)进行微调,用于序列标注任务。
  • 采用严格和宽松的F1得分评估模型性能,以衡量在提取16项特征时的精确率与召回率。
  • 另开展了一项关联任务,以评估模型在多结节报告中将提取的特征正确关联至对应结节的能力。
  • 选择GatorTron(一种在超过900亿词语料上预训练的大语言模型)因其在提取和关联任务中均表现出色。

实验结果

研究问题

  • RQ1哪些基于Transformer的NLP模型在从超声报告中提取16种临床相关的甲状腺结节特征方面表现最佳?
  • RQ2在该临床NLP任务中,像GatorTron这样的大规模语言模型相较于BERT和RoBERTa等较小模型的性能如何?
  • RQ3标准化的标注指南是否能提升从放射科报告中标注甲状腺结节特征的一致性与可靠性?
  • RQ4NLP模型在描述多个结节的报告中,多大程度上能准确地将提取的特征关联到正确的甲状腺结节?
  • RQ5该框架在支持基于电子健康记录的甲状腺结节管理结果研究和文档质量评估方面具有多大潜力?

主要发现

  • GatorTron在从超声报告中提取16项甲状腺结节特征方面,严格F1得分达到0.8851,宽松F1得分达到0.9495。
  • GatorTron在提取和关联任务中均优于所有其他模型,包括BERT、RoBERTa、LongFormer和DeBERTa。
  • 该模型在复杂报告中表现出强大的泛化能力,在关联任务中实现了0.9321的F1得分,成功将特征与正确结节关联。
  • 本研究建立了标准化的标注框架,可在多样化的超声报告中实现对临床相关特征的一致标注。
  • 这是首次系统性地将基于Transformer的NLP应用于从非结构化临床文本中提取大量、临床全面的甲状腺结节特征。
  • 研究结果为自动化评估超声报告质量以及基于电子健康记录的甲状腺结节管理结果研究奠定了基础。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。