Skip to main content
QUICK REVIEW

[Paper Review] Predicting Climate Variability over the Indian Region Using Data Mining Strategies

M. Naresh Kumar|arXiv (Cornell University)|Sep 23, 2015
Climate variability and models5 references3 citations
TL;DR

This paper proposes a data mining framework combining Expectation Maximization (EM) clustering and Support Vector Machine (SVM) regression to predict climate variability over India. It outperforms k-means and linear regression, achieving an RMSE of 1.19 in the Montane region and 0.89 in the Humid Subtropical region, while identifying a new climate zone in the Himalayan foothills not captured by traditional Koppen classification.

ABSTRACT

In this paper an approach based on expectation maximization (EM) clustering to find the climate regions and a support vector machine to build a predictive model for each of these regions is proposed. To minimize the biases in the estimations a ten cross fold validation is adopted both for obtaining clusters and building the predictive models. The EM clustering could identify all the zones as per the Koppen classification over Indian region. The proposed strategy when employed for predicting temperature has resulted in an RMSE of $1.19$ in the Montane climate region and $0.89$ in the Humid Sub Tropical region as compared to $2.9$ and $0.95$ respectively predicted using k-means and linear regression method.

Motivation & Objective

  • To improve climate variability prediction in India by developing a robust regional modeling approach using data mining.
  • To identify climate regions from long-term gridded climate data using EM clustering, avoiding initial cluster bias.
  • To compare the proposed EM-SVM framework with traditional k-means and linear regression in terms of prediction accuracy.
  • To validate the model using 10-fold cross-validation and assess performance via RMSE across distinct climate zones.
  • To investigate climate change impacts by identifying new climate regions not aligned with standard Koppen classifications.

Proposed method

  • Apply EM clustering to long-term (1948–2012) 2.5°×2.5° gridded climate data to identify homogeneous climate regions.
  • Use a mixture of Gaussian distributions to model cluster centroids, iteratively optimizing means, covariances, and prior probabilities.
  • Perform 10-fold cross-validation to ensure robust estimation of RMSE and reduce overfitting in both clustering and modeling stages.
  • Build a separate SVM regression model for each climate region using historical climate variables as predictors.
  • Train models on data from j−p years and validate predictions for the next p years using actual observed values.
  • Compute RMSE between predicted and observed temperature values at each grid point to evaluate model performance.

Experimental results

Research questions

  • RQ1Can EM clustering effectively identify climate regions in India that align with or improve upon the Koppen classification?
  • RQ2How does the EM-SVM framework compare to k-means and linear regression in predicting regional temperature variability?
  • RQ3What is the impact of data dimensionality and numerical precision on EM clustering performance in high-dimensional climate datasets?
  • RQ4Does the proposed method detect emerging climate zones due to climate change, such as a new region in the Himalayan foothills?
  • RQ5To what extent does 10-fold cross-validation reduce bias in RMSE estimation for climate prediction models?

Key findings

  • EM clustering successfully identified seven climate regions, including a new region comprising Uttarakhand, Sikkim, and Arunachal Pradesh, distinct from the traditional Montane zone.
  • The proposed EM-SVM model achieved an RMSE of 1.19 in the Montane region, significantly lower than the 2.9 RMSE from k-means and linear regression.
  • In the Humid Subtropical region, the EM-SVM model achieved an RMSE of 0.89, outperforming the 0.95 RMSE from k-means and linear regression.
  • The EM-SVM framework showed superior performance in the Montane and Humid Subtropical zones, but performance degradation was observed with increasing data dimensionality.
  • The spatial maps of predicted 2012 temperatures and absolute errors confirmed the model’s accuracy, with lower error concentrations in stable climate zones.
  • The results suggest that climate change may be altering regional climate patterns, as evidenced by the emergence of a new climatic cluster in the Himalayan region.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.