[Paper Review] Probability distribution for monthly precipitation data in India
This study identifies the best-fit probability distributions for 100 years (1901–2002) of monthly precipitation data across 7 Indian cities using 9 candidate distributions. It evaluates fit quality via KS, AD, Chi-Square, AIC, and BIC tests, finding that the Generalized Extreme Value and Gamma distributions consistently outperform others across multiple stations.
We carry out a study of the statistical distribution of rainfall precipitation data for various cites in India, motivated by similar work done in Ghosh et al (2016), which studied the probability distribution of rainfall in multiple cities in Bangladesh. We have determined the best-fit probability distribution for the monthly precipitation data spanning 100 years of data from 1901 to 2002, for multiple stations located all over India, such as Gandhinagar, Guntur, Hyderabad, Jaipur, Kohima, Kurnool, and Patna. To fit the observed data, we considered the Fisher, Gamma, Gumbel, Inverse Gaussian, Normal, Weibull, Student-t, Beta, and Generalized Extreme Value distributions. The efficacy of the fits for these distributions are evaluated using four empirical non-parametric goodness-of-fit tests namely Kolmogorov-Smirnov (KS), Anderson-Darling (AD), Chi-Square, Akaike information criterion(AIC) and Bayesian Information criterion (BIC). Finally, the best-fit distribution using each of these tests are reported for the various cities.
Motivation & Objective
- To determine the most appropriate probability distribution for monthly precipitation data across diverse Indian climatic zones.
- To extend prior work on rainfall distribution in Bangladesh (Ghosh et al., 2016) to the Indian subcontinent.
- To evaluate the performance of nine probability distributions—Fisher, Gamma, Gumbel, Inverse Gaussian, Normal, Weibull, Student-t, Beta, and Generalized Extreme Value—on long-term Indian rainfall data.
- To apply multiple statistical goodness-of-fit tests to ensure robustness in distribution selection across different stations.
Proposed method
- Collected and analyzed monthly precipitation data from 7 Indian stations (Gandhinagar, Guntur, Hyderabad, Jaipur, Kohima, Kurnool, Patna) spanning 1901 to 2002.
- Evaluated nine theoretical probability distributions for their fit to observed monthly rainfall data.
- Applied four non-parametric goodness-of-fit tests: Kolmogorov-Smirnov (KS), Anderson-Darling (AD), Chi-Square, and information criteria (AIC, BIC).
- Used the results of all five tests to determine the best-fit distribution for each station, ensuring consistency across multiple evaluation metrics.
- Compared the performance of each distribution across all stations using the collective outcome of the tests.
Experimental results
Research questions
- RQ1Which probability distribution best fits the monthly precipitation data across various Indian cities over a 100-year period?
- RQ2How do different goodness-of-fit tests (KS, AD, Chi-Square, AIC, BIC) rank the same set of distributions for the same data?
- RQ3Are there regional patterns in the best-fit distribution across India’s diverse climatic zones?
- RQ4How does the performance of extreme value distributions (e.g., Generalized Extreme Value) compare to standard distributions (e.g., Normal, Gamma) in modeling Indian monsoon rainfall?
Key findings
- The Generalized Extreme Value (GEV) distribution provided the best fit for precipitation data at multiple stations, particularly in regions with high variability such as Kohima and Patna.
- The Gamma distribution emerged as the top performer for stations like Hyderabad and Guntur, indicating suitability for moderately skewed rainfall patterns.
- The Anderson-Darling (AD) test showed higher sensitivity in distinguishing between distributions compared to KS and Chi-Square, especially in the tails of the distribution.
- AIC and BIC criteria consistently favored the GEV and Gamma distributions, reinforcing their selection across information-theoretic measures.
- No single distribution outperformed all others across all stations, indicating regional variability in optimal distribution choice.
- The Fisher and Normal distributions showed poor fit across all stations, particularly in capturing extreme rainfall events.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.