Skip to main content
QUICK REVIEW

[Paper Review] Nonparametric approaches for analyzing carbon emission: from statistical and machine learning perspectives

Yiming Ma, Hang Liu|arXiv (Cornell University)|Mar 27, 2023
Environmental Impact and SustainabilityEnvironmental Science3 citations
TL;DR

This study evaluates nonparametric methods—kernel regression, random forest, and neural networks—for modeling urban carbon emissions using data from ten Chinese cities (2005–2019). It finds that neural networks outperform linear models and other nonparametric techniques in fitting and predicting carbon emissions, offering a more accurate tool for policy-relevant forecasting and factor analysis, particularly for Wuhu city, where emissions peak around 2025 under targeted interventions.

ABSTRACT

Linear regression models, especially the extended STIRPAT model, are routinely-applied for analyzing carbon emissions data. However, since the relationship between carbon emissions and the influencing factors is complex, fitting a simple parametric model may not be an ideal solution. This paper investigated various nonparametric approaches in statistics and machine learning (ML) for modeling carbon emissions data, including kernel regression, random forest and neural network. We selected data from ten Chinese cities from 2005 to 2019 for modeling studies. We found that neural network had the best performance in both fitting and prediction accuracy, which implies its capability of expressing the complex relationships between carbon emissions and the influencing factors. This study provides a new means for quantitative modeling of carbon emissions research that helps to understand how to characterize urban carbon emissions and to propose policy recommendations for "carbon reduction". In addition, we used the carbon emissions data of Wuhu city as an example to illustrate how to use this new approach.

Motivation & Objective

  • To address the limitations of linear STIRPAT models in capturing complex, nonlinear relationships between carbon emissions and influencing factors.
  • To evaluate the performance of nonparametric statistical and machine learning methods in modeling urban carbon emissions.
  • To provide a data-driven, high-accuracy modeling framework for forecasting carbon emissions and informing climate policy.
  • To demonstrate the application of neural networks in analyzing and predicting emissions trends using city-level data, with Wuhu as a case study.

Proposed method

  • The study applies the extended STIRPAT model with population, affluence (GDP per capita), energy intensity, and industrial structure (share of secondary sector) as independent variables.
  • It compares four modeling approaches: linear log-linear regression, kernel regression, random forest, and feedforward neural networks on panel data from ten Chinese cities (2005–2019).
  • Model performance is evaluated using in-sample fit (R²) and out-of-sample prediction accuracy (MSE, MAE) across the ten cities.
  • A neural network model is selected for detailed analysis and forecasting in Wuhu, using scenario-based projections with varying assumptions on energy intensity, GDP growth, and industrial structure.
  • The model incorporates feature importance analysis to assess the relative impact of each factor on emissions.
  • Two scenarios are projected: a baseline (no policy intervention) and a policy intervention scenario emphasizing industrial upgrading and energy efficiency.

Experimental results

Research questions

  • RQ1How do nonparametric models compare to traditional linear STIRPAT models in fitting and predicting urban carbon emissions?
  • RQ2Which nonparametric method—kernel regression, random forest, or neural network—delivers the highest predictive accuracy for carbon emissions in Chinese cities?
  • RQ3What are the relative contributions of population, affluence, energy intensity, and industrial structure to carbon emissions in a high-dimensional, nonlinear context?
  • RQ4How can nonparametric modeling inform policy interventions for achieving carbon peak and neutrality in urban settings?

Key findings

  • Neural networks achieved the best fit and prediction accuracy among all models tested, with significantly lower mean squared error (MSE) and mean absolute error (MAE) compared to linear regression and other nonparametric methods.
  • The neural network model demonstrated superior ability to capture complex, nonlinear interactions between carbon emissions and influencing factors such as energy intensity and industrial structure.
  • In Wuhu, under a baseline scenario with no policy intervention, carbon emissions are projected to continue growing, with no peak by 2030.
  • Under a policy intervention scenario involving a 4% annual decrease in energy intensity, 3% GDP per capita growth, and 2% annual decline in secondary industry share, emissions peak around 2025 and begin to decline.
  • The study confirms that optimizing industrial structure and reducing energy intensity are critical levers for controlling emissions growth, even at the cost of moderate economic growth.
  • Feature importance analysis from the neural network model indicates that energy intensity and industrial structure are the most influential factors, followed by affluence and population.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.