[论文解读] Electronic health record phenotyping improves detection and screening of type 2 diabetes in the general United States population: A cross-sectional, unselected, retrospective study
本研究证明,通过利用多变量逻辑回归和随机森林模型,基于电子健康记录(EHR)的表型分析可显著提高美国普通人群中2型糖尿病的检出率。完整EHR模型在性能上显著优于仅使用体质指数(BMI)、年龄、性别和生活方式因素的传统风险模型(p<0.001),并发现了新的关联,如偏头痛和心律失常与2型糖尿病呈负相关。
Objectives: In the United States, 25% of people with type 2 diabetes are undiagnosed. Conventional screening models use limited demographic information to assess risk. We evaluated whether electronic health record (EHR) phenotyping could improve diabetes screening, even when records are incomplete and data are not recorded systematically across patients and practice locations. Methods: In this cross-sectional, retrospective study, data from 9,948 US patients between 2009 and 2012 were used to develop a pre-screening tool to predict current type 2 diabetes, using multivariate logistic regression. We compared (1) a full EHR model containing prescribed medications, diagnoses, and traditional predictive information, (2) a restricted EHR model where medication information was removed, and (3) a conventional model containing only traditional predictive information (BMI, age, gender, hypertensive and smoking status). We additionally used a random-forests classification model to judge whether including additional EHR information could increase the ability to detect patients with Type 2 diabetes on new patient samples. Results: Using a patient's full or restricted EHR to detect diabetes was superior to using basic covariates alone (p<0.001). The random forests model replicated on out-of-bag data. Migraines and cardiac dysrhythmias were negatively associated with type 2 diabetes, while acute bronchitis and herpes zoster were positively associated, among other factors. Conclusions: EHR phenotyping resulted in markedly superior detection of type 2 diabetes in a general US population, could increase the efficiency and accuracy of disease screening, and are capable of picking up signals in real-world records.
研究动机与目标
- 提高美国人群中未诊断2型糖尿病的检出率,目前仍有25%的患者未被确诊。
- 评估尽管存在数据记录不完整或不一致的情况,基于EHR的表型是否能提升筛查准确性。
- 比较基于EHR的模型与仅依赖人口统计学和基本临床因素的传统风险模型的性能表现。
- 利用真实世界EHR数据,识别与2型糖尿病相关的新型临床关联。
- 通过随机森林模型的样本外预测,验证EHR表型分析的稳健性。
提出的方法
- 对2009至2012年期间来自9,948名美国患者的EHR数据进行横断面、回顾性分析。
- 基于完整的EHR数据(包括诊断、药物使用和传统预测因子)构建多变量逻辑回归模型,以预测当前2型糖尿病状态。
- 构建一个排除药物数据的受限EHR模型,以评估其对预测性能的影响。
- 构建一个仅使用基本协变量(BMI、年龄、性别、高血压和吸烟状态)的传统模型。
- 应用随机森林分类模型,基于袋外数据评估预测性能,并识别新的信号关联。
- 使用受试者工作特征(ROC)分析,比较所有三个模型的性能表现。
实验结果
研究问题
- RQ1与传统筛查模型相比,EHR表型分析是否能显著提升2型糖尿病的检出率?
- RQ2药物数据的纳入在多大程度上影响基于EHR的糖尿病筛查模型的预测准确性?
- RQ3在真实世界EHR数据中,除传统危险因素外,哪些新型临床关联可预测2型糖尿病?
- RQ4当患者和医疗机构间数据记录不完整或不一致时,EHR表型分析是否仍能保持高性能?
- RQ5随机森林模型在EHR表型分析中,能在多大程度上复现并推广逻辑回归模型的发现?
主要发现
- 完整EHR模型在检测2型糖尿病方面显著优于传统模型(p < 0.001),表现出更优的判别能力。
- 即使在排除药物数据的受限EHR模型中,其性能仍显著优于传统模型(p < 0.001),表明仅凭诊断和临床数据即可提升检出率。
- 随机森林模型在袋外数据上成功复现了结果,证实了模型的稳健性和泛化能力。
- 偏头痛和心律失常与2型糖尿病呈负相关,提示可能存在保护性或相反的临床信号。
- 急性支气管炎和带状疱疹与2型糖尿病呈正相关,提示可能存在共病关系或预测关联。
- EHR表型分析能有效从真实世界中未结构化且不完整的临床记录中提取有意义的生物信号,从而提升筛查效率和准确性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。