[论文解读] Synthetic Photovoltaic and Wind Power Forecasting Data
本文提出一个大规模、公开可用的合成数据集,涵盖德国境内120个光伏电站和273个风电站的发电预测,基于真实气象观测数据和物理模型生成。该数据集支持可再生能源预测领域的真实机器学习研究,结果显示合成数据上的模型误差与真实历史数据高度一致,为未来研究提供了可靠的基准。
Photovoltaic and wind power forecasts in power systems with a high share of renewable energy are essential in several applications. These include stable grid operation, profitable power trading, and forward-looking system planning. However, there is a lack of publicly available datasets for research on machine learning based prediction methods. This paper provides an openly accessible time series dataset with realistic synthetic power data. Other publicly and non-publicly available datasets often lack precise geographic coordinates, timestamps, or static power plant information, e.g., to protect business secrets. On the opposite, this dataset provides these. The dataset comprises 120 photovoltaic and 273 wind power plants with distinct sides all over Germany from 500 days in hourly resolution. This large number of available sides allows forecasting experiments to include spatial correlations and run experiments in transfer and multi-task learning. It includes side-specific, power source-dependent, non-synthetic input features from the ICON-EU weather model. A simulation of virtual power plants with physical models and actual meteorological measurements provides realistic synthetic power measurement time series. These time series correspond to the power output of virtual power plants at the location of the respective weather measurements. Since the synthetic time series are based exclusively on weather measurements, possible errors in the weather forecast are comparable to those in actual power data. In addition to the data description, we evaluate the quality of weather-prediction-based power forecasts by comparing simplified physical models and a machine learning model. This experiment shows that forecasts errors on the synthetic power data are comparable to real-world historical power measurements.
研究动机与目标
- 为解决基于机器学习的可再生能源预测领域缺乏公开、全面数据集的问题。
- 提供大规模、逼真的合成数据集,包含精确的地理坐标、时间戳和电站物理特性。
- 通过为每个电站提供详细的静态元数据,支持迁移学习、多任务学习和零样本学习等高级机器学习研究。
- 通过在合成数据上对比物理模型与梯度提升回归树(GBRT)模型,建立预测误差的基准。
- 通过基于实际数值天气预报(NWP)输入生成数据,确保合成数据真实反映现实世界中的预测误差。
提出的方法
- 利用基于真实历史气象观测数据(来自Icosahedral Nonhydrostatic,ICON模型)的物理模型生成合成功率时间序列。
- 数据集包含500天的小时分辨率数据(2018年12月8日至2020年6月2日),覆盖全年四季,支持模型的稳健训练与测试。
- 每个电站均分配独特的物理与几何参数,如光伏组件倾角与朝向、风力涡轮机转子-发电机比等,以确保真实感并支持迁移学习。
- 非合成输入特征(如风速、全球水平辐照度和温度)源自ICON-NWP模型,并与合成功率输出配对。
- 训练并对比梯度提升回归树(GBRT)模型与物理基线模型(如Enercon和McLean功率曲线),以评估预测精度。
- 通过支持截断训练集(如7至365天)的实验,使数据集适用于数据稀缺条件下的评估。
实验结果
研究问题
- RQ1机器学习模型在合成数据上的预测误差与在真实历史数据上的误差相比如何?
- RQ2合成数据能否可靠支持可再生能源发电预测模型的基准测试,特别是在迁移学习和多任务学习中?
- RQ3在不同量级的训练数据下,物理模型与机器学习模型(如GBRT)之间的性能差距如何?
- RQ4当缺乏涡轮机特定参数时,不同经验功率曲线(如McLean、Enercon)的准确性如何比较?
- RQ5详细电站专属元数据的引入在多大程度上提升了模型泛化能力与迁移学习表现?
主要发现
- GBRT与物理基线模型在合成数据集上的预测误差,与真实世界历史功率测量中的误差高度一致。
- 在光伏预测中,当有足够的训练数据(如365天)时,GBRT模型优于物理基线模型,平均nRMSE为0.072,而物理模型为0.085。
- 在风力发电预测中,GBRT模型始终优于Enercon物理基线模型,全量训练数据下平均nRMSE从基线的0.210降至0.125。
- 在训练数据有限(如7天)时,物理模型对光伏预测更具鲁棒性,而GBRT模型即使仅有14天数据,对风力预测的表现也更优。
- 当缺乏涡轮机特定参数时,McLean经验功率曲线可作为风力预测的可行基线,其nRMSE值接近Enercon基线(0.212–0.239)。
- 由于包含电站专属元数据(如朝向、转子直径和轮毂高度),该合成数据集可支持可靠地评估迁移学习与零样本学习。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。