[论文解读] Biases in human mobility data impact epidemic modeling
本文识别并量化了人类移动数据中的两种主要偏差——技术可及性偏差和此前被忽视的数据生成偏差,其中高财富个体产生的移动数据 disproportionately(不成比例地)更多,导致流行病模型出现偏差。通过一种重新采样按财富五分位组并校正覆盖范围的去偏框架,作者表明,有偏数据会导致SIR模拟中疾病传播速度和严重程度被高估,呼吁研究人员在基于移动数据的建模中考虑数据公平性。
Large-scale human mobility data is a key resource in data-driven policy making and across many scientific fields. Most recently, mobility data was extensively used during the COVID-19 pandemic to study the effects of governmental policies and to inform epidemic models. Large-scale mobility is often measured using digital tools such as mobile phones. However, it remains an open question how truthfully these digital proxies represent the actual travel behavior of the general population. Here, we examine mobility datasets from multiple countries and identify two fundamentally different types of bias caused by unequal access to, and unequal usage of mobile phones. We introduce the concept of data generation bias, a previously overlooked type of bias, which is present when the amount of data that an individual produces influences their representation in the dataset. We find evidence for data generation bias in all examined datasets in that high-wealth individuals are overrepresented, with the richest 20% contributing over 50% of all recorded trips, substantially skewing the datasets. This inequality is consequential, as we find mobility patterns of different wealth groups to be structurally different, where the mobility networks of high-wealth users are denser and contain more long-range connections. To mitigate the skew, we present a framework to debias data and show how simple techniques can be used to increase representativeness. Using our approach we show how biases can severely impact outcomes of dynamic processes such as epidemic simulations, where biased data incorrectly estimates the severity and speed of disease transmission. Overall, we show that a failure to account for biases can have detrimental effects on the results of studies and urge researchers and practitioners to account for data-fairness in all future studies of human mobility.
研究动机与目标
- 探究大规模人类移动数据中的偏差(尤其是来自移动电话CDR的数据)对流行病建模结果的影响。
- 识别并表征一种新型偏差,称为“数据生成偏差”,即富裕个体因使用频率更高而产生更多移动数据。
- 开发并验证一种去偏框架,以在社会经济群体之间恢复移动网络的代表性。
- 评估有偏与去偏移动数据对动态流行病模拟的影响,特别是在传播速度和严重程度方面的差异。
- 倡导在流行病响应等政策相关应用中,重视移动数据使用的公平性。
提出的方法
- 本研究分析了塞拉利昂、刚果民主共和国和伊拉克的通话详单(CDR)数据集,以检测不同财富五分位组之间的移动数据差异。
- 提出一种基于重采样的去偏方法:通过从原始流量矩阵中使用多项式抽样,将行程在财富五分位组之间均等重新分配,以纠正数据生成偏差。
- 通过按观测频率成比例抽样流量,该方法保留了网络结构,确保空间连通性模式保持现实。
- 通过使用用户渗透率对移动网络覆盖不全的地区中的流量进行重新缩放,来校正技术可及性偏差。
- 对于无数据的地区,使用基于最贫穷五分位组移动模式拟合的重力模型来估算缺失流量,从而保持网络密度。
- 在原始网络和去偏网络上,使用R₀ = 2.5和μ = 1/6天的SIR元群模型运行流行病模拟,以比较结果。
实验结果
研究问题
- RQ1由不同财富群体间数据生成不均所引发的数据生成偏差,在多大程度上影响了移动数据集的代表性?
- RQ2高财富与低财富用户群体之间的移动网络结构差异(如连通性、长程连接)如何变化?
- RQ3有偏移动数据对流行病模拟中疾病传播速度和严重程度的预测有何影响?
- RQ4基于重采样的去偏框架是否能有效恢复移动网络的代表性,同时不扭曲空间模式?
- RQ5对技术可及性偏差和数据生成偏差的联合校正,在多大程度上提升了流行病建模的准确性?
主要发现
- 最富裕的20%用户在所有分析数据集中贡献了超过50%的记录行程,表明存在强烈的数据生成偏差。
- 高财富群体的移动网络显著更密集,且包含更多长程连接,相较于低财富群体。
- 去偏后,流行病模拟中疾病传播速度相比使用有偏数据的模拟降低了20%至30%。
- 在使用去偏数据的模拟中,流行病高峰被延迟,整体攻击率更低,表明有偏数据高估了传播潜力。
- 基于重力模型的缺失流量插补方法在低覆盖区域提升了网络完整性,同时保持了现实的空间结构。
- 本研究表明,若不校正数据生成偏差,将系统性地导致流行病预测失真,尤其在低收入和中等收入国家中更为显著。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。