[Paper Review] A new regression model for positive data
This paper proposes a novel regression model for positive continuous data using a new parameterization of the beta prime (BP) distribution based on mean and precision parameters. The model features a quadratic variance function, enables separate regression structures for mean and precision, and offers improved flexibility for skewed data compared to exponential family models, with estimation via maximum likelihood and diagnostic tools for influence analysis.
In this paper, we propose a regression model where the response variable is beta prime distributed using a new parameterization of this distribution that is indexed by mean and precision parameters. The proposed regression model is useful for situations where the variable of interest is continuous and restricted to the positive real line and is related to other variables through the mean and precision parameters. The variance function of the proposed model has a quadratic form. In addition, the beta prime model has properties that its competitor distributions of the exponential family do not have. Estimation is performed by maximum likelihood. Furthermore, we discuss residuals and influence diagnostic tools. Finally, we also carry out an application to real data that demonstrates the usefulness of the proposed model.
Motivation & Objective
- To develop a flexible regression model for positive continuous response variables that are skewed and bounded away from zero.
- To address limitations of generalized linear models (GLMs) when data exhibit high skewness or non-linear variance structures.
- To propose a new parameterization of the beta prime distribution using mean and precision parameters for improved interpretability and modeling flexibility.
- To enable separate regression structures for both the mean and precision parameters, enhancing model adaptability.
- To provide robust inference tools, including maximum likelihood estimation, residuals, and local influence diagnostics for model validation.
Proposed method
- Proposes a new parameterization of the beta prime (BP) distribution using the mean $\mu$ and precision $\phi$ parameters instead of shape parameters $\alpha$ and $\beta$.
- Derives the probability density function (PDF) of the BP distribution in terms of $\mu$ and $\phi$, enabling direct modeling of the mean and precision.
- Constructs a regression model where the response variable $Y_i$ follows a BP distribution with mean $\mu_i$ and precision $\phi_i$, both linked to covariates via link functions $g_1$ and $g_2$.
- Employs maximum likelihood estimation (MLE) for parameter inference, with iterative algorithms to solve the likelihood equations.
- Develops residuals and local influence diagnostics to assess model fit and detect influential observations.
- Applies perturbation schemes (regressor, precision covariate, and simultaneous perturbations) to derive influence measures using the perturbation matrix $\widehat{\mathbf{\Delta}}$.
Experimental results
Research questions
- RQ1Can a new parameterization of the beta prime distribution improve modeling flexibility for positive continuous data with skewness?
- RQ2How does the proposed regression model with mean and precision regression structures compare to standard GLMs in fitting skewed positive data?
- RQ3What is the impact of different perturbation schemes on the influence of covariates in the model?
- RQ4How effective are the proposed residuals and local influence diagnostics in detecting influential observations?
- RQ5Does the quadratic variance function in the BP model provide better fit than standard variance functions in exponential family models?
Key findings
- The proposed model exhibits a quadratic variance function of the form $\text{Var}(Y_i) = \mu_i^2(1 + \mu_i)\phi_i^{-1}$, which captures complex heteroscedasticity patterns not available in standard GLMs.
- The new parameterization in terms of mean and precision improves interpretability and enables separate modeling of location and dispersion.
- Maximum likelihood estimation is feasible and effective for the proposed model, with convergence achieved through iterative algorithms.
- Diagnostic tools, including residuals and local influence measures, are successfully derived and applied to detect influential observations.
- The application to real data demonstrates superior fit and flexibility compared to competing models, especially for highly skewed positive data.
- The perturbation-based influence analysis using $\widehat{\mathbf{\Delta}}$ matrices effectively identifies influential cases under various perturbation schemes.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.