[论文解读] Beta and Kumaraswamy distributions as non-nested hypotheses in the modeling of continuous bounded data
本文提出了一种基于似然比的筛选准则,用于在连续有界数据中比较贝塔分布与柯马克-瓦米分布,利用对数似然比统计量的渐近分布来估计正确选择概率(PCS)。该方法选择PCS较高的模型,通过纳入模型选择的不确定性,优于仅依赖最大对数似然值的赤池信息准则(AIC),在模拟和真实数据应用中均表现出优异的准确性。
Nowadays, beta and Kumaraswamy distributions are the most popular models to fit continuous bounded data. These models present some characteristics in common and to select one of them in a practical situation can be of great interest. With this in mind, in this paper we propose a method of selection between the beta and Kumaraswamy distributions. We use the logarithm of the likelihood ratio statistic (denoted by $T_n$, where $n$ is the sample size) and obtain its asymptotic distribution under the hypotheses $H_{\mathcal B}$ and $H_{\mathcal K}$, where $H_{\mathcal B}$ ($H_{\mathcal K}$) denotes that the data come from the beta (Kumaraswamy) distribution. Since both models has the same number of parameters, based on the Akaike criterion, we choose the model that has the greater log-likelihood value. We here propose to use the probability of correct selection (given by $P(T_n>0)$ or $P(T_n<0)$ depending on the null hypothesis) instead of only to observe the maximized log-likelihood values. We obtain an approximation for the probability of correct selection under the hypotheses $H_{\mathcal B}$ and $H_{\mathcal K}$ and select the model that maximizes it. A simulation study is presented in order to evaluate the accuracy of the approximated probabilities of correct selection. We illustrate our method of selection in two applications to real data sets involving proportions.
研究动机与目标
- 为解决在建模连续有界数据时,缺乏正式方法以区分贝塔分布与柯马克-瓦米分布的问题。
- 改进仅依赖最大对数似然值而未考虑选择不确定性的赤池信息准则(AIC)。
- 在非嵌套假设下,基于正确选择概率(PCS)开发一种选择准则。
- 推导在贝塔分布与柯马克-瓦米分布原假设下,PCS的渐近近似。
- 通过模拟研究和涉及比例的真实世界数据应用,验证所提方法。
提出的方法
- 该方法使用贝塔模型与柯马克-瓦米模型之间最大对数似然值的比值的对数作为检验统计量 $T_n$。
- 在 $H_{\text{B}}$(数据来自贝塔分布)和 $H_{\text{K}}$(数据来自柯马克-瓦米分布)下,推导出 $T_n$ 的渐近正态性,从而实现对 PCS 的估计。
- 利用在每种原假设下 $T_n$ 的渐近均值(AM)和渐近方差(AV)对正确选择概率进行近似。
- 选择在 $H_{\text{B}}$ 下 $P(T_n > 0)$ 或在 $H_{\text{K}}$ 下 $P(T_n < 0)$ 的估计值更高的模型。
- 该方法采用 AM 和 AV 的渐近近似来计算 PCS,无需重采样,从而提升计算效率。
- 通过模拟研究验证该方法,并将其应用于两个真实数据集:穆斯林人口比例和无神论者人口比例。
实验结果
研究问题
- RQ1能否利用渐近理论可靠地近似非嵌套的贝塔分布与柯马克-瓦米分布的正确选择概率(PCS)?
- RQ2所提出的基于 PCS 的选择准则是否在选择有界连续数据的正确模型方面优于赤池信息准则(AIC)?
- RQ3在不同样本大小和参数值下,贝塔分布与柯马克-瓦米分布原假设下的 PCS 渐近近似精度如何?
- RQ4在模型区分中,达到预设正确选择概率所需的最小样本量是多少?
- RQ5在涉及比例数据(如全球宗教比例)的真实世界应用中,该方法表现如何?
主要发现
- 在穆斯林人口比例数据集中,贝塔模型的估计 PCS 为 0.7174,柯马克-瓦米模型为 0.5917,因此选择贝塔分布。
- 在无神论者人口比例数据集中,柯马克-瓦米模型的估计 PCS 为 0.7872,贝塔模型为 0.6812,因此选择柯马克-瓦米分布。
- 模拟研究证实,PCS 的渐近近似与经验概率非常接近,各种参数设置和样本大小下的平均相对误差低于 5%。
- 该方法在两种原假设下均达到 95% 正确选择概率的最小样本量为 200,具体取决于参数值。
- 在模拟中,基于 PCS 的选择准则始终优于 AIC,因为它通过概率推理考虑了模型选择的不确定性。
- 模拟得到的经验 PCS 值与渐近估计值接近(例如,第一个数据集中贝塔分布的 0.7370 对比 0.7174),验证了近似的准确性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。