[论文解读] Social media mining for identification and exploration of health-related information from pregnant women
本研究提出了一种混合NLP与机器学习的流程,用于识别社交媒体上的孕妇,收集其纵向健康发帖,并分析孕期各季度的药物使用模式。通过基于规则的检测与有监督分类器(F1值0.81),从Twitter和DailyStrength平台识别出34,895名用户的发帖时间线,尽管提及频率较低,仍能检测到可识别的药物摄入模式。
Widespread use of social media has led to the generation of substantial amounts of information about individuals, including health-related information. Social media provides the opportunity to study health-related information about selected population groups who may be of interest for a particular study. In this paper, we explore the possibility of utilizing social media to perform targeted data collection and analysis from a particular population group -- pregnant women. We hypothesize that we can use social media to identify cohorts of pregnant women and follow them over time to analyze crucial health-related information. To identify potentially pregnant women, we employ simple rule-based searches that attempt to detect pregnancy announcements with moderate precision. To further filter out false positives and noise, we employ a supervised classifier using a small number of hand-annotated data. We then collect their posts over time to create longitudinal health timelines and attempt to divide the timelines into different pregnancy trimesters. Finally, we assess the usefulness of the timelines by performing a preliminary analysis to estimate drug intake patterns of our cohort at different trimesters. Our rule-based cohort identification technique collected 53,820 users over thirty months from Twitter. Our pregnancy announcement classification technique achieved an F-measure of 0.81 for the pregnancy class, resulting in 34,895 user timelines. Analysis of the timelines revealed that pertinent health-related information, such as drug-intake and adverse reactions can be mined from the data. Our approach to using user timelines in this fashion has produced very encouraging results and can be employed for other important tasks where cohorts, for which health-related information may not be available from other sources, are required to be followed over time to derive population-based estimates.
研究动机与目标
- 为解决孕期女性缺乏纵向健康数据的问题,特别是那些未被纳入上市前临床试验的群体。
- 开发一种可扩展的方法,用于识别并追踪Twitter和DailyStrength等社交媒体平台上的孕妇。
- 从用户时间线中提取并分析随时间推移的健康相关信息,特别是药物使用情况及不良反应。
- 评估将社交媒体作为脆弱人群药物流行病学监测的补充数据源的可行性。
- 建立一种可推广的流程,用于在代表性不足的人群群体中进行队列识别与纵向健康监测。
提出的方法
- 采用基于规则的关键词搜索,检测社交媒体发帖中关于怀孕的声明,重点关注如“pregnant”、“due date”和“baby bump”等术语。
- 应用有监督的二分类器(基于100条人工标注的推文训练)从初始基于规则的结果中过滤假阳性,对怀孕类别的F1值达到0.81。
- 从已识别的用户处随时间收集纵向发帖,构建Twitter和DailyStrength平台上的个人健康时间线。
- 使用时间启发式方法与语言线索将发帖分配至特定孕期季度,但该方法被指出存在不足。
- 通过包含常用名称、缩写及拼写变体,扩展药物提及检测,以提高后处理阶段的召回率。
- 结合多种NLP技术,包括命名实体识别与时间关系抽取,对用户时间线中的健康事件与药物暴露进行分类。
实验结果
研究问题
- RQ1社交媒体数据能否有效用于识别和追踪孕妇,以实现纵向健康监测?
- RQ2基于规则与机器学习的技术在社交媒体文本中检测怀孕声明的准确性如何?
- RQ3在多大程度上能可靠地从用户时间线中提取与健康相关的信息,如药物使用与不良反应?
- RQ4药物使用模式在社交媒体发帖中如何随孕期各季度变化?
- RQ5所提出的流程能否推广至其他代表性不足或难以接触的人群群体,用于公共卫生研究?
主要发现
- 基于规则的队列识别方法在30个月内成功收集了53,820名用户,其中34,895名经分类器确认为孕妇。
- 怀孕声明分类器的F1值达到0.81,表明在识别社交媒体内容中自我报告的怀孕方面具有高精确率与高召回率。
- 来自34,895名用户的纵向时间线使药物摄入模式得以检测,其中布洛芬在第一、第二和第三孕期分别被76、72和90名用户提及。
- 尽管提及频率较低(如大多数药物低于0.5%),各孕期仍观察到一致的药物使用模式,表明存在可检测的趋势。
- 该方法成功从DailyStrength平台检索到11,435名用户的257,531条发帖,更长的发帖内容有助于更详细地提取健康信息。
- 局限性包括对时态与时间顺序处理不佳、推文中性别指代模糊,以及孕期分类不够理想,需在未来进一步优化。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。