[論文レビュー] Beta and Kumaraswamy distributions as non-nested hypotheses in the modeling of continuous bounded data
本稿では、連続的かつ有界なデータのモデリングにおいて、ベータ分布と Kumaraswamy 分布の間で尤度比に基づく選択基準を提案する。尤度比統計量の漸近的分布を用いて、正しい選択の確率(PCS)を推定し、より高い PCS を示すモデルを選択する。この手法は、モデル選択の不確実性を組み込むことで、AIC より優れた性能を示し、シミュレーションおよび実データ応用において高い正確性を示している。
Nowadays, beta and Kumaraswamy distributions are the most popular models to fit continuous bounded data. These models present some characteristics in common and to select one of them in a practical situation can be of great interest. With this in mind, in this paper we propose a method of selection between the beta and Kumaraswamy distributions. We use the logarithm of the likelihood ratio statistic (denoted by $T_n$, where $n$ is the sample size) and obtain its asymptotic distribution under the hypotheses $H_{\mathcal B}$ and $H_{\mathcal K}$, where $H_{\mathcal B}$ ($H_{\mathcal K}$) denotes that the data come from the beta (Kumaraswamy) distribution. Since both models has the same number of parameters, based on the Akaike criterion, we choose the model that has the greater log-likelihood value. We here propose to use the probability of correct selection (given by $P(T_n>0)$ or $P(T_n<0)$ depending on the null hypothesis) instead of only to observe the maximized log-likelihood values. We obtain an approximation for the probability of correct selection under the hypotheses $H_{\mathcal B}$ and $H_{\mathcal K}$ and select the model that maximizes it. A simulation study is presented in order to evaluate the accuracy of the approximated probabilities of correct selection. We illustrate our method of selection in two applications to real data sets involving proportions.
研究の動機と目的
- 連続的かつ有界なデータのモデリングにおいて、ベータ分布と Kumaraswamy 分布を区別する正式な手法の欠如に応えること。
- 最大対数尤度値のみに依存するが、選択の不確実性を考慮しない AIC(赤赤情報量基準)の改善を図ること。
- 非ネストされた仮説のもとで、正しい選択の確率(PCS)に基づく選択基準の開発。
- ベータ分布および Kumaraswamy 分布の帰無仮説のもとでの PCS の漸近的近似を導出すること。
- シミュレーション研究および割合を含む実世界のデータ応用を通じて、提案手法の妥当性を検証すること。
提案手法
- 本手法は、ベータ分布モデルと Kumaraswamy 分布モデルの間の最大対数尤度の比の対数を統計量 $T_n$ として用いる。
- 帰無仮説 $H_{ ext{B}}$(データがベータ分布に従う)および $H_{ ext{K}}$(データが Kumaraswamy 分布に従う)の下で $T_n$ の漸近的正規性が導かれる。これにより PCS の推定が可能になる。
- 各仮説の下での $T_n$ の漸近的平均(AM)および漸近的分散(AV)を用いて、正しい選択の確率を近似する。
- $H_{ ext{B}}$ の下で $P(T_n > 0)$、$H_{ ext{K}}$ の下で $P(T_n < 0)$ を計算し、より高い推定 PCS を示すモデルを選択する。
- 再サンプリングに依存せずに、AM および AV の漸近的近似を用いて PCS を計算することで、計算効率を向上させる。
- シミュレーション研究および 2 つの実データセット(ムスリム人口の割合および無神論者人口の割合)への応用を通じて、手法の妥当性を検証する。
実験結果
リサーチクエスチョン
- RQ1非ネストされたベータ分布および Kumaraswamy 分布に対して、漸近的理論を用いて正しい選択の確率(PCS)を信頼性高く近似できるか?
- RQ2提案された PCS に基づく選択基準は、有界な連続的データの正しいモデル選択において、AIC より優れた性能を示すか?
- RQ3さまざまなサンプルサイズおよびパラメータ値の下で、ベータ分布および Kumaraswamy 分布の帰無仮説のもとでの PCS の漸近的近似はどの程度正確か?
- RQ4モデル識別において、事前に定めた正しい選択確率に到達するための最小サンプルサイズは何か?
- RQ5グローバルな宗教的割合のような実世界の割合データを含む応用において、本手法はどの程度の性能を示すか?
主な発見
- ムスリム人口の割合データセットでは、ベータ分布の推定 PCS が 0.7174、Kumaraswamy 分布の推定 PCS が 0.5917 であり、ベータ分布が選択された。
- 無神論者人口の割合データセットでは、Kumaraswamy 分布の推定 PCS が 0.7872、ベータ分布の推定 PCS が 0.6812 であり、Kumaraswamy 分布が選択された。
- シミュレーション研究により、PCS の漸近的近似が実証確率とよく一致することが確認され、さまざまなパラメータ設定およびサンプルサイズで平均相対誤差が 5% 未満であった。
- 両仮説のもとで 95% の正しい選択確率に到達するための最小サンプルサイズは、パラメータ値に応じて 200 であった。
- シミュレーションにおいて、PCS に基づく選択基準は AIC より一貫して優れており、確率的推論によるモデル選択の不確実性を考慮しているためである。
- シミュレーションからの実証 PCS 値は、漸近的推定値と近い値であった(例:最初のデータセットにおけるベータ分布で 0.7370 対 0.7174)、近似の正確性が裏付けられた。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。