[论文解读] Fine-tuning foundational models to code diagnoses from veterinary health records
本研究在科罗拉多州立大学兽医教学医院的246,473份人工编码的兽医电子健康记录(EHR)笔记上微调了十个预训练的大规模语言模型(LLMs),以实现对7,739种不同诊断代码的SNOMED-CT诊断编码自动化。使用大型、临床专用LLM并结合大量标注数据可获得最佳性能,表明自动化编码可显著提升互操作性,并通过支持跨物种数据整合,促进整体健康(One Health)研究。
Veterinary medical records represent a large data resource for application to veterinary and One Health clinical research efforts. Use of the data is limited by interoperability challenges including inconsistent data formats and data siloing. Clinical coding using standardized medical terminologies enhances the quality of medical records and facilitates their interoperability with veterinary and human health records from other sites. Previous studies, such as DeepTag and VetTag, evaluated the application of Natural Language Processing (NLP) to automate veterinary diagnosis coding, employing long short-term memory (LSTM) and transformer models to infer a subset of Systemized Nomenclature of Medicine - Clinical Terms (SNOMED-CT) diagnosis codes from free-text clinical notes. This study expands on these efforts by incorporating all 7,739 distinct SNOMED-CT diagnosis codes recognized by the Colorado State University (CSU) Veterinary Teaching Hospital (VTH) and by leveraging the increasing availability of pre-trained language models (LMs). 13 freely-available pre-trained LMs were fine-tuned on the free-text notes from 246,473 manually-coded veterinary patient visits included in the CSU VTH's electronic health records (EHRs), which resulted in superior performance relative to previous efforts. The most accurate results were obtained when expansive labeled data were used to fine-tune relatively large clinical LMs, but the study also showed that comparable results can be obtained using more limited resources and non-clinical LMs. The results of this study contribute to the improvement of the quality of veterinary EHRs by investigating accessible methods for automated coding and support both animal and human health research by paving the way for more integrated and comprehensive health databases that span species and institutions.
研究动机与目标
- 为解决兽医医疗记录不一致和数据孤立的问题,实现自动化、标准化的临床编码。
- 通过与SNOMED-CT等标准化术语集成,提升兽医电子健康记录(EHR)的互操作性。
- 通过通用编码标准实现人类与动物健康记录之间的数据关联,支持整体健康研究。
- 评估多种预训练LLM在从自由文本临床笔记中诊断兽医疾病方面的性能。
- 识别在资源有限环境中实现可扩展、高精度诊断编码的最优模型与数据配置。
提出的方法
- 在科罗拉多州立大学兽医教学医院(CSU VTH)EHR系统提供的246,473份人工编码的兽医就诊记录上,微调了十个可免费获取的预训练LLM,包括GatorTron、MedicalAI ClinicalBERT、VetBERT和GPT-2变体。
- 采用多类别分类框架,将自由文本临床笔记映射到7,739个不同的SNOMED-CT诊断代码。
- 利用OMOP通用数据模型(CDM)对编码进行标准化,以实现未来跨机构和跨物种的数据整合。
- 采用迁移学习技术,将通用和临床专用LLM适配至兽医诊断编码任务。
- 使用标准自然语言处理(NLP)指标(如F1分数、精确率和召回率)在多标签分类任务中评估模型性能。
- 探讨了在实现高编码准确率时,模型规模、临床专业化程度与数据可用性之间的权衡。
实验结果
研究问题
- RQ1预训练的大规模语言模型能否被有效微调,以实现兽医EHR中SNOMED-CT诊断编码的自动化?
- RQ2模型架构与临床专业化程度如何影响对7,739种兽医诊断代码的编码准确率?
- RQ3在有限兽医数据上微调时,非临床LLM在多大程度上可达到与临床专用模型相当的性能?
- RQ4训练数据规模对LLM在兽医诊断编码中性能的影响如何?
- RQ5微调后的LLM能否通过标准化的SNOMED-CT编码,实现兽医与人类健康记录之间的互操作性?
主要发现
- 通过在246,473次兽医就诊的大量标注数据上微调大型、临床专用LLM(如MedicalAI ClinicalBERT和VetBERT),实现了最精确的诊断编码。
- 微调GatorTron及其他大型通用LLM也取得了高性能,表明无需专用兽医模型即可实现最先进的结果。
- 当在足够数据上进行微调时,较小的非临床LLM(如BERT和RoBERTa)也获得了可比的编码准确率,表明其在资源有限环境中的可扩展性。
- 与先前方法(如DeepTag和VetTag)相比,本研究在多标签F1分数上实现了显著提升,后者仅限于SNOMED-CT代码的子集。
- 使用OMOP CDM和标准化SNOMED-CT编码,可实现未来人类与动物健康系统之间的数据关联,支持整体健康倡议。
- 结果表明,大规模、高精度的自动化诊断编码是可行的,可显著提升兽医临床研究中的数据质量和互操作性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。