[论文解读] Predicting drug recalls from Internet search engine queries
本研究提出利用Bing搜索引擎的聚合互联网搜索查询量来预测即将发生的美国食品药品监督管理局(FDA)药品召回。通过分析特定药品在各州的查询量激增情况,该方法在提前一天预测召回时实现了0.791的AUC值和约6倍的5%截断处提升,且对处方药和中等风险药品的预测性能更高。
Batches of pharmaceutical are sometimes recalled from the market when a safety issue or a defect is detected in specific production runs of a drug. Such problems are usually detected when patients or healthcare providers report abnormalities to medical authorities. Here we test the hypothesis that defective production lots can be detected earlier by monitoring queries to Internet search engines. We extracted queries from the USA to the Bing search engine which mentioned one of 5,195 pharmaceutical drugs during 2015 and all recall notifications issued by the Food and Drug Administration (FDA) during that year. By using attributes that quantify the change in query volume at the state level, we attempted to predict if a recall of a specific drug will be ordered by FDA in a time horizon ranging from one to 40 days in future. Our results show that future drug recalls can indeed be identified with an AUC of 0.791 and a lift at 5% of approximately 6 when predicting a recall will occur one day ahead. This performance degrades as prediction is made for longer periods ahead. The most indicative attributes for prediction are sudden spikes in query volume about a specific medicine in each state. Recalls of prescription drugs and those estimated to be of medium-risk are more likely to be identified using search query data. These findings suggest that aggregated Internet search engine data can be used to facilitate in early warning of faulty batches of medicines.
研究动机与目标
- 探究互联网搜索查询模式是否可作为FDA药品召回的早期信号。
- 确定特定药品搜索量的突然增加是否与后续召回相关。
- 评估搜索查询数据在预测未来40天内召回事件中的预测性能。
- 确定哪些药品类别(如处方药、风险等级)最适用于通过搜索数据进行预测。
- 探索利用公开的搜索引擎数据构建实时药品安全问题早期预警系统的可行性。
提出的方法
- 收集了2015年美国各州Bing搜索引擎中5,195种药品的每日搜索查询量。
- 从FDA的公开数据库中提取了同期的召回通知信息。
- 基于查询量的变化构建特征,特别关注各州层面上的突然激增。
- 使用这些查询量变化特征训练机器学习分类器,以预测未来1至40天内是否会发生召回。
- 采用AUC和前5%预测结果的精确率(5%处提升)评估模型性能。
- 按药品类型(如处方药与非处方药)和召回风险等级(低、中、高)进行子组分析。
实验结果
研究问题
- RQ1特定药品的互联网搜索量突然增加是否可预测即将发生的FDA药品召回?
- RQ2使用搜索查询数据预测药品召回的准确性如何?性能是否随预测时间范围变化?
- RQ3哪些类型的药品(如处方药、非处方药)最能通过搜索数据进行预测?
- RQ4召回风险等级(低、中、高)是否影响通过搜索查询模式检测召回的能力?
- RQ5各州层面上的搜索量变化是否能提升召回预测的及时性和准确性?
主要发现
- 当预测提前一天时,模型的受试者工作特征曲线下面积(AUC)达到0.791。
- 模型在5%处的提升约为6,表明其在前5%的召回事件中识别精度是随机选择的六倍。
- 随着预测时间范围的延长,预测性能下降,提前40天预测的AUC值更低。
- 特定药品在州层面的搜索量突然激增是最具指示性的预测特征。
- 处方药和中等风险药品更可能通过搜索查询数据被提前检测到。
- 结果表明,聚合的搜索引擎数据可作为缺陷药品批次的可行早期预警信号。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。