Skip to main content
QUICK REVIEW

[論文レビュー] Statistical Analysis of Data Repeatability Measures

Zeyi Wang, Eric Bridgeford|arXiv (Cornell University)|May 25, 2020
Statistical Methods and Inference参考文献 45被引用数 11
ひとこと要約

本稿は、一変量および多変量(M)ANOVAモデルにおけるさまざまな分布仮定下で、データ再現性の統計的パワーを評価する。具体的には、判別能(Disc)、ランク和検定、F検定、ICC推定値を対象とする。非正規性や測定回数の増加下でも、Discはランク和検定およびICCに基づく手法を常に上回る性能を示す。一方、ファングリープ法は最も性能が低く、対称性の破れはすべての手法の性能を低下させる。

ABSTRACT

The advent of modern data collection and processing techniques has seen the size, scale, and complexity of data grow exponentially. A seminal step in leveraging these rich datasets for downstream inference is understanding the characteristics of the data which are repeatable -- the aspects of the data that are able to be identified under a duplicated analysis. Conflictingly, the utility of traditional repeatability measures, such as the intraclass correlation coefficient, under these settings is limited. In recent work, novel data repeatability measures have been introduced in the context where a set of subjects are measured twice or more, including: fingerprinting, rank sums, and generalizations of the intraclass correlation coefficient. However, the relationships between, and the best practices among these measures remains largely unknown. In this manuscript, we formalize a novel repeatability measure, discriminability. We show that it is deterministically linked with the correlation coefficient under univariate random effect models, and has desired property of optimal accuracy for inferential tasks using multivariate measurements. Additionally, we overview and systematically compare repeatability statistics using both theoretical results and simulations. We show that the rank sum statistic is deterministically linked to a consistent estimator of discriminability. The power of permutation tests derived from these measures are compared numerically under Gaussian and non-Gaussian settings, with and without simulated batch effects. Motivated by both theoretical and empirical results, we provide methodological recommendations for each benchmark setting to serve as a resource for future analyses. We believe these recommendations will play an important role towards improving repeatability in fields such as functional magnetic resonance imaging, genomics, pharmacology, and more.

研究の動機と目的

  • さまざまな分布仮定下で、データ再現性を測定するための判別能(Disc)、ランク和検定、F検定、ICC推定値の相対的統計的パワーを評価すること。
  • 非正規性(例:対数正規誤差)およびバッチ効果が再現性測定の性能に与える影響を調査すること。
  • 繰り返し測定回数の増加が、さまざまな統計的検定のパワーに与える影響を評価すること。
  • 被験者間の対称性が破れた場合に、これらの測定法のロバストネスを検討すること。
  • ファングリープベースの手法を他の再現性指標と比較し、統計的パワーの観点から評価すること。

提案手法

  • 被験者固有のランダム効果と独立誤差項を含む一変量および多変量(M)ANOVAモデルをシミュレートする。
  • 10,000回のモンテカルロシミュレーションを用いて、サンプルサイズ(n = 5 から 100)の変動に伴う第一種過誤率および統計的パワーを推定する。
  • 被験者間分散と被験者内分散の比に基づき、判別能(Disc)を再現性の指標として用いる。
  • DiscおよびF検定と比較するため、パーミュテーション検定およびノンパラメトリックなランク和検定を適用する。
  • ICCを被験者間分散と全分散の比として定義し、多変量設定ではトレースに基づくICCを用いる。
  • 分布仮定の破れを評価するため、非ガウス誤差構造(例:対数正規)を導入する。

実験結果

リサーチクエスチョン

  • RQ1一変量および多変量設定下で、判別能(Disc)はランク和検定や距離ベースのウィルコクソン検定よりも高い統計的パワーを有するか?
  • RQ2非正規誤差分布下で、DiscのパワーはF検定およびICC推定値と比較してどうなるか?
  • RQ3繰り返し測定回数の増加が、Discとランク和検定またはF検定の相対的パフォーマンスに与える影響は何か?
  • RQ4バッチ効果が、ガウス(M)ANOVAモデルにおけるDisc、ランク和検定、F検定のパワーに与える影響は何か?
  • RQ5ファングリープベースの再現性推定は、どのような条件下で他の手法よりも性能が低くなるか?

主な発見

  • 非正規性下でも、Discは一変量および多変量設定において、ランク和検定およびICCに基づく手法を一貫して上回る性能を示す。
  • 非ガウス(対数正規)誤差モデル下でも、Discはランク和検定およびICC推定値よりも高いパワーを維持するが、ICCの決定的変換はもはや成り立たない。
  • 繰り返し測定回数の増加に伴い、Discの他の手法に対するパワー優位性は拡大する。複数のランク和を組み合わせても、この優位性は完全には消失しない。
  • バッチ効果下では、ガウスモデルにおいてランク和検定がDiscを上回る性能を示す。これは文脈依存的な優位性を示す。
  • ファングリープベースの手法は、すべてのシミュレーション設定(一変量、多変量、非スパースケース)において一貫して最もパワーが低い。
  • 被験者間の対称性が破れた場合、すべての手法がパワーを失う。これは、このような条件下での再現性評価に根本的な制限があることを示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。