[Paper Review] Wind energy forecasting with missing values within a fully conditional specification framework
This paper proposes a fully conditional specification (FCS)-based universal imputation framework for wind power forecasting that jointly handles missing input features and target variables during both model estimation and operational forecasting. By treating forecasting as a joint imputation and prediction problem under missing-at-random assumptions, the method outperforms the standard 'impute-then-predict' approach, especially in probabilistic forecasting, while maintaining consistency across stages and reducing overfitting risks.
Wind power forecasting is essential to power system operation and electricity markets. As abundant data became available thanks to the deployment of measurement infrastructures and the democratization of meteorological modelling, extensive data-driven approaches have been developed within both point and probabilistic forecasting frameworks. These models usually assume that the dataset at hand is complete and overlook missing value issues that often occur in practice. In contrast to that common approach, we rigorously consider here the wind power forecasting problem in the presence of missing values, by jointly accommodating imputation and forecasting tasks. Our approach allows inferring the joint distribution of input features and target variables at the model estimation stage based on incomplete observations only. We place emphasis on a fully conditional specification method owing to its desirable properties, e.g., being assumption-free when it comes to these joint distributions. Then, at the operational forecasting stage, with available features at hand, one can issue forecasts by implicitly imputing all missing entries. The approach is applicable to both point and probabilistic forecasting, while yielding competitive forecast quality within both simulation and real-world case studies. It confirms that by using a powerful universal imputation method like fully conditional specification, the proposed approach is superior to the common approach, especially in the context of probabilistic forecasting.
Motivation & Objective
- To address the widespread issue of missing values in wind power forecasting datasets, which commonly leads to data loss and forecast degradation.
- To develop a consistent forecasting framework that jointly models imputation and prediction, avoiding inconsistencies between model estimation and operational forecasting stages.
- To evaluate the performance of the proposed method in both point and probabilistic forecasting settings under varying missing data rates.
- To compare the FCS-based approach with conventional 'impute-then-predict' strategies, particularly in terms of forecast accuracy and robustness.
- To assess computational feasibility and scalability of the method as dimensionality increases.
Proposed method
- The method employs multiple imputation via fully conditional specification (FCS), which models the conditional distribution of each variable given all others in a sequential, iterative manner.
- At model estimation, parameters are estimated using only observed data under the missing-at-random (MAR) assumption, without requiring complete cases.
- At operational forecasting, both input features and target variables are treated as missing and iteratively imputed using the learned conditional models.
- The approach naturally supports probabilistic forecasting by generating multiple imputed realizations from the joint distribution of inputs and targets.
- The framework is applicable to both point and probabilistic forecasting and is compatible with various base models, including quantile regression and deep learning.
- The method avoids the inconsistency of 'impute-then-predict' by ensuring that imputation during inference aligns with the joint modeling used during training.
Experimental results
Research questions
- RQ1How does the proposed FCS-based universal imputation framework compare to the standard 'impute-then-predict' approach in terms of forecast accuracy under missing data?
- RQ2Can the joint modeling of imputation and forecasting improve probabilistic forecast quality compared to sequential imputation and prediction?
- RQ3To what extent does the method mitigate overfitting in the presence of missing data?
- RQ4How does forecast performance degrade with increasing missing data rates, and how does the FCS method compare under such conditions?
- RQ5What are the computational costs of the FCS-based approach, and can it be scaled to high-dimensional wind power datasets?
Key findings
- The FCS-based method outperforms the 'impute-then-predict' approach in both point and probabilistic forecasting, with a more pronounced advantage in probabilistic settings.
- The method reduces overfitting risks by sharing information across variables and maintaining consistency between model estimation and operational forecasting stages.
- Forecast quality degrades with increasing missing data rates, but the FCS approach maintains superior performance across all tested rates.
- The training time of the FCS model is higher than that of standard quantile regression (QR) models but remains manageable compared to deep learning baselines like DeepAR.
- Operational inference time for the FCS method is low (0.01 seconds per forecast), making it suitable for real-time applications.
- The approach enables consistent probabilistic forecasting through multiple imputed realizations, naturally capturing uncertainty in both inputs and targets.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.