[论文解读] A Double Machine Learning Trend Model for Citizen Science Data
本文提出了一种双机器学习趋势模型,用于从公民科学数据中估计物种种群趋势,解决了因采样协议不一致导致的年度间混杂问题。通过利用机器学习估计倾向得分,并采用基于模拟的方法校正残余混杂,该模型利用eBird数据在27公里分辨率下生成了准确、空间明确的趋势估计,方向和变化幅度的准确性均较高。
1. Citizen and community-science (CS) datasets have great potential for estimating interannual patterns of population change given the large volumes of data collected globally every year. Yet, the flexible protocols that enable many CS projects to collect large volumes of data typically lack the structure necessary to keep consistent sampling across years. This leads to interannual confounding, as changes to the observation process over time are confounded with changes in species population sizes. 2. Here we describe a novel modeling approach designed to estimate species population trends while controlling for the interannual confounding common in citizen science data. The approach is based on Double Machine Learning, a statistical framework that uses machine learning methods to estimate population change and the propensity scores used to adjust for confounding discovered in the data. Additionally, we develop a simulation method to identify and adjust for residual confounding missed by the propensity scores. Using this new method, we can produce spatially detailed trend estimates from citizen science data. 3. To illustrate the approach, we estimated species trends using data from the CS project eBird. We used a simulation study to assess the ability of the method to estimate spatially varying trends in the face of real-world confounding. Results showed that the trend estimates distinguished between spatially constant and spatially varying trends at a 27km resolution. There were low error rates on the estimated direction of population change (increasing/decreasing) and high correlations on the estimated magnitude. 4. The ability to estimate spatially explicit trends while accounting for confounding in citizen science data has the potential to fill important information gaps, helping to estimate population trends for species, regions, or seasons without rigorous monitoring data.
研究动机与目标
- 解决由于不同年份采样协议不一致导致的公民科学数据中的年度间混杂问题。
- 开发一种稳健的统计模型,用于在调整可观测和不可观测混杂因素的同时估计物种种群趋势。
- 使在缺乏严格长期监测数据的区域或物种种群中实现空间详细的趋势估计成为可能。
- 在现实混杂条件下验证该方法在区分空间恒定与空间可变种群趋势方面的表现。
提出的方法
- 该方法采用双机器学习技术,在通过机器学习模型估计的倾向得分基础上调整混杂因素,以估计种群趋势。
- 采用两阶段机器学习框架:首先为每个观测值估计倾向得分,以建模其在某一年被纳入的概率;然后在这些得分的条件下估计结果(例如,物种种群的检测情况)。
- 引入一种基于模拟的方法,以检测并校正倾向得分未能捕捉的残余混杂因素,从而提高模型的稳健性。
- 该模型应用于eBird数据,实现了在地理区域范围内27公里分辨率的空间显式趋势估计。
- 该框架将因果推断原理与现代机器学习技术相结合,以应对公民科学中常见的高维、非随机采样模式。
实验结果
研究问题
- RQ1在公民科学数据中存在因采样不一致导致的年度间混杂时,双机器学习方法能否准确估计物种种群趋势?
- RQ2在真实世界条件下,该方法在区分空间恒定与空间可变种群趋势方面表现如何?
- RQ3与标准双机器学习相比,基于模拟的残余混杂校正方法在多大程度上提升了趋势估计的准确性?
- RQ4该模型在不同地理区域中估计种群变化方向和幅度方面的表现如何?
主要发现
- 在模拟研究中,该模型成功在27公里空间分辨率下区分了空间恒定与空间可变的种群趋势。
- 估计种群变化方向(上升或下降)的错误率较低,表明趋势方向检测具有高度可靠性。
- 估计趋势幅度与真实趋势幅度之间的相关性较高,表明在量化种群变化速率方面具有很强的准确性。
- 基于模拟的残余混杂校正方法有效提升了模型性能,降低了未观测混杂因素带来的偏差。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。