[論文レビュー] A Robust Spearman Correlation Coefficient Permutation Test
本稿では、正規性が満たされない場合や小標本サイズ下でも従来の検定がType Iエラーを適切に制御できない問題を解決するため、スケール化された検定静標(studentized test statistic)に基づく頑健な順列検定を提案する。この手法は、さまざまな分布的仮定と標本サイズにおいて正確なエラー率を維持し、シミュレーションおよびゲノム・臨床データ応用において標準的手法を上回る性能を示す。
In this work, we show that Spearman's correlation coefficient test about $H_0:ρ_s=0$ found in most statistical software packages is theoretically incorrect and performs poorly when bivariate normality assumptions are not met or the sample size is small. The historical works about these tests make an unverifiable assumption that the approximate bivariate normality of original data justifies using classic approximations. In general, there is common misconception that the tests about $ρ_s=0$ are robust to deviations from bivariate normality. In fact, we found under certain scenarios violation of the bivariate normality assumption has severe effects on type I error control for the most commonly utilized tests. To address this issue, we developed a robust permutation test for testing the general hypothesis $H_0: ρ_s=0$. The proposed test is based on an appropriately studentized statistic. We will show that the test is theoretically asymptotically valid in the general setting when two paired variables are uncorrelated but dependent. This desired property was demonstrated across a range of distributional assumptions and sample sizes in simulation studies, where the proposed test exhibits robust type I error control across a variety of settings, even when the sample size is small. We demonstrated the application of this test in real world examples of transcriptomic data of the TCGA breast cancer patients and a data set of PSA levels and age.
研究の動機と目的
- 二変量正規性の仮定が満たされない場合、あるいは標本サイズが小さい場合に従来のスピアマン相関検定がType Iエラーを適切に制御できない問題を解決すること。
- 一般の条件下(相関のない従属性を含む)で理論的に妥当かつ実用的に頑健な順列検定を開発すること。
- 正規性仮定が検証不能な標準的手法(t検定、フィッシャーのZ変換、単純な順列検定)の信頼できる代替手法を提供すること。
- 正規性仮定が頻繁に破られるが実世界の応用(トランスクリプトームデータや臨床バイオマーカー研究)において、正確な推論を保証すること。
提案手法
- スケール化されたスピアマン相関係数に基づく順列検定を提案し、頑健性を向上させる。
- 検定静標 $ t = r_s \sqrt{\frac{n-2}{1 - r_s^2}} $ を使用するが、分散の安定化を図るため、順列をスケール化後に帰無仮説下の分布に適用する。
- 帰無仮説における交換可能性の下で順列を適用するが、小標本での精度向上のため、スケール化された検定静標を用いる。
- さまざまな非正規分布と標本サイズにおけるType Iエラー制御を評価するために、シミュレーション研究を実施する。
- TCGA乳癌コhortデータおよびPSA値と年齢の関係を用いた実データを用いて、従来の検定と比較して手法を検証する。
- 片側対立仮説 $ H_1: \rho_s > 0 $ に対応するため、PSAデータに負の対数変換を適用して対立仮説と整合させる。
実験結果
リサーチクエスチョン
- RQ1二変量正規性の仮定が満たされない場合、スピアマン相関の従来のt検定は適切なType Iエラー制御を維持するか?
- RQ2スケール化された静標に基づく順列検定は、小標本および非正規性下でもType Iエラー制御を改善できるか?
- RQ3提案手法は、フィッシャーのZ、単純な順列検定、漸近正規近似と比較して、エラー率制御においてどのように異なるか?
- RQ4提案手法は、標本サイズが増加しても二変量正規性からの逸脱に対して頑健か?
- RQ5変数が相関のないが従属である場合でも、この手法は有効性を保つのか?
主な発見
- 従来のt検定は、非正規性下でも大標本サイズであっても、Type Iエラー率が著しく上昇する。
- 提案されたスケール化順列検定は、すべてのシミュレーションされた分布と標本サイズ(n=10を含む)で正確なType Iエラー制御を維持する。
- TCGA乳癌データでは、提案手法のp値は0.081であったが、従来手法はp < 0.05を報告しており、本手法がより信頼できる結果を示している。
- PSAデータでは、すべての検定が帰無仮説を棄却(p < 0.001)したが、提案手法は非正規な周辺分布および二変量分布下でも一貫した性能を示した。
- フィッシャー=ヤーツ変換は、非正規性下でもType Iエラーの上昇を是正できず、周辺変換の限界を示している。
- 4次のモーメントが有限であっても、本手法は頑健であり、独立性に限定されない一般の従属構造下でも正しいエラー率に収束する。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。