Skip to main content
QUICK REVIEW

[论文解读] Urban Air Pollution Forecasting: a Machine Learning Approach leveraging Satellite Observations and Meteorological Forecasts

Giacomo Blanco, Luca Barco|arXiv (Cornell University)|May 30, 2024
Air Quality Monitoring and Forecasting被引用 4
一句话总结

本研究提出了一种机器学习框架,利用Sentinel-5P卫星数据、气象变量和地形特征预测城市空气污染,在米兰对五种主要污染物的平均绝对百分比误差(MAPE)约为30%。该模型在无需依赖地面监测站的情况下进行训练,使其适用于数据匮乏的城市地区。

ABSTRACT

Air pollution poses a significant threat to public health and well-being, particularly in urban areas. This study introduces a series of machine-learning models that integrate data from the Sentinel-5P satellite, meteorological conditions, and topological characteristics to forecast future levels of five major pollutants. The investigation delineates the process of data collection, detailing the combination of diverse data sources utilized in the study. Through experiments conducted in the Milan metropolitan area, the models demonstrate their efficacy in predicting pollutant levels for the forthcoming day, achieving a percentage error of around 30%. The proposed models are advantageous as they are independent of monitoring stations, facilitating their use in areas without existing infrastructure. Additionally, we have released the collected dataset to the public, aiming to stimulate further research in this field. This research contributes to advancing our understanding of urban air quality dynamics and emphasizes the importance of amalgamating satellite, meteorological, and topographical data to develop robust pollution forecasting models.

研究动机与目标

  • 解决缺乏地面监测基础设施的城市中城市空气污染预测的挑战。
  • 通过整合多源数据(卫星观测(Sentinel-5P)、气象预报和地形特征)提高预测准确性。
  • 开发一种可扩展的、与监测站无关的预测系统,通过基于网格的预测方法适用于各类城市区域。
  • 通过可靠的提前一天污染预测,支持主动的公共卫生和政策干预。
  • 通过公开发布一个全面的城市空气质量数据集,激发进一步的研究。

提出的方法

  • 在融合数据集上训练多种机器学习模型——XGBoost、SGD回归和线性回归,该数据集结合了卫星反演的污染物浓度、气象变量和地形特征。
  • 采用时间窗口方法,使用1天、7天和14天的历史数据窗口,以捕捉短期和中期污染趋势。
  • 通过年度交叉验证,每次保留一年数据用于验证,其余年份用于训练。
  • 引入空间网格化(500米分辨率),将预测结果外推至整个城市区域,超越监测站的位置。
  • 利用多时相卫星观测,捕捉污染物浓度随时间的动态变化。
  • 通过标准回归指标(MAE、MAPE和RMSE)优化模型性能,评估所有污染物和窗口大小下的表现。

实验结果

研究问题

  • RQ1仅使用卫星、气象和地形数据而无需依赖地面监测站,机器学习模型能否有效预测城市空气污染?
  • RQ2与传统方法相比,引入多时相卫星数据和气象变量如何提升预测准确性?
  • RQ3在预测次日污染水平时,1天、7天或14天的最优历史时间窗口是什么?该选择在不同污染物之间是否存在差异?
  • RQ4在多种污染物和窗口设置下,不同机器学习算法(XGBoost、SGD、线性回归)的性能表现如何比较?
  • RQ5通过空间网格化,该模型在多大程度上能实现跨城市区域的泛化,从而实现在无监测站区域的预测?

主要发现

  • XGBoost模型在所有污染物和评估指标下均表现最佳,始终优于SGD和线性回归。
  • 所有五种污染物(PM2.5、PM10、NO2、O3、SO2)的平均绝对百分比误差(MAPE)在所有时间窗口下均约为30%。
  • 对于PM10,更长的时间窗口(14天)提升了预测准确性;而对于PM2.5和NO2,七天窗口效果最佳。
  • O3和SO2的预测在一天时间窗口下最为准确,表明这些污染物具有独特的时序动态特征。
  • 尽管MAE值较高,SGD回归在一天和七天窗口下对SO2的MAPE表现更优,表明其相对误差性能更佳。
  • 该模型在城市网格上的外推预测中表现出稳健性,使无监测站区域的污染制图成为可能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。