Skip to main content
QUICK REVIEW

[论文解读] Tracking Air Pollution in China: Near Real-Time PM2.5 Retrievals from Multiple Data Sources

Guannan Geng, Qingyang Xiao|arXiv (Cornell University)|Mar 11, 2021
Air Quality and Health Impacts参考文献 45被引用 6
一句话总结

本文提出了中国空气质量追踪数据库(TAP),该数据库为2000年至今的中国提供了近实时、高分辨率(10公里)的PM2.5数据集,通过整合地面监测数据、卫星AOD数据、排放清单和化学传输模型输出,采用两阶段机器学习模型结合SMOTE与基于树的插补方法生成。模型的袋外R²达到0.83,有效捕捉了高污染事件和缺失的AOD数据,支持及时的空气质量监测与长期环境分析。

ABSTRACT

Air pollution has altered the Earth radiation balance, disturbed the ecosystem and increased human morbidity and mortality. Accordingly, a full-coverage high-resolution air pollutant dataset with timely updates and historical long-term records is essential to support both research and environmental management. Here, for the first time, we develop a near real-time air pollutant database known as Tracking Air Pollution in China (TAP, tapdata.org) that combines information from multiple data sources, including ground measurements, satellite retrievals, dynamically updated emission inventories, operational chemical transport model simulations and other ancillary data. Daily full-coverage PM2.5 data at a spatial resolution of 10 km is our first near real-time product. The TAP PM2.5 is estimated based on a two-stage machine learning model coupled with the synthetic minority oversampling technique and a tree-based gap-filling method. Our model has an averaged out-of-bag cross-validation R2 of 0.83 for different years, which is comparable to those of other studies, but improves its performance at high pollution levels and fills the gaps in missing AOD on daily scale. The full coverage and near real-time updates of the daily PM2.5 data allow us to track the day-to-day variations in PM2.5 concentrations over China in a timely manner. The long-term records of PM2.5 data since 2000 will also support policy assessments and health impact studies. The TAP PM2.5 data are publicly available through our website for sharing with the research and policy communities.

研究动机与目标

  • 开发覆盖完整、高分辨率且近实时的中国PM2.5数据集,以支持环境研究与政策制定。
  • 通过融合多种数据源,解决地面监测稀疏与卫星AOD数据不完整的问题。
  • 利用先进的机器学习技术,提升PM2.5估算精度,特别是在高污染事件期间。
  • 为健康影响与政策评估研究提供自2000年以来的长期历史记录。
  • 通过持续的数据更新,实现实时追踪中国各地每日PM2.5变化。

提出的方法

  • TAP PM2.5数据集通过两阶段机器学习模型生成,该模型融合了地面PM2.5观测、卫星AOD反演、动态排放清单和化学传输模型输出。
  • 模型采用合成少数类过采样技术(SMOTE),以提升在训练数据中代表性不足的高污染事件中的表现。
  • 采用基于树的插补方法,重建缺失的每日AOD数据,提升数据完整性。
  • 模型结合地面观测与辅助数据(包括气象与土地利用变量)进行训练。
  • 通过袋外R²进行交叉验证,评估模型在不同年份的表现。
  • 最终产品提供10公里空间分辨率、每日更新的全境PM2.5地图,具备近实时性。

实验结果

研究问题

  • RQ1如何利用多种异构数据源可靠地估算中国范围内近实时、高分辨率的PM2.5数据?
  • RQ2卫星AOD、地面观测与排放清单的融合在多大程度上提升了PM2.5估算精度,特别是在污染事件期间?
  • RQ3结合SMOTE与插补技术的机器学习模型能否有效处理缺失的AOD数据,并在高污染条件下保持精度?
  • RQ4模型在不同年份与不同污染水平下的表现如何变化?其预测鲁棒性如何?
  • RQ5该数据集在多大程度上可支持长期环境与健康影响评估?

主要发现

  • TAP模型在不同年份的袋外交叉验证R²达到0.83,表明其具有强劲的预测性能。
  • 模型在高污染条件下的精度显著提升,这对健康与政策应用至关重要。
  • 合成少数类过采样技术(SMOTE)有效增强了模型对罕见但高影响污染事件的处理能力。
  • 基于树的插补方法成功重建了缺失的每日AOD数据,提升了数据完整性和模型可靠性。
  • 该数据集自2000年以来提供10公里分辨率、全境覆盖的每日PM2.5估算,支持长期趋势分析。
  • TAP PM2.5数据集通过专用网站公开发布,供研究人员与政策制定者使用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。