[论文解读] QR-Adjustment for Clustering Tests Based on Nearest Neighbor Contingency Tables
本文提出对最近邻列联表(NNCT)检验采用QR校正,以校正完全空间随机性(CSR)独立性下的条件推理问题,使用经验估计的共享(Q)和反射(R)最近邻期望值。蒙特卡洛模拟显示,QR校正对经验大小或统计功效无显著影响,表明在聚类检测中,未校正与校正后的检验结果相当。
The spatial interaction between two or more classes of points may cause spatial clustering patterns such as segregation or association, which can be tested using a nearest neighbor contingency table (NNCT). A NNCT is constructed using the frequencies of class types of points in nearest neighbor (NN) pairs. For the NNCT-tests, the null pattern is either complete spatial randomness (CSR) of the points from two or more classes (called CSR independence) or random labeling (RL). The distributions of the NNCT-test statistics depend on the number of reflexive NNs (denoted by $R$) and the number of shared NNs (denoted by $Q$), both of which depend on the allocation of the points. Hence $Q$ and $R$ are fixed quantities under RL, but random variables under CSR independence. Using their observed values in NNCT analysis makes the distributions of the NNCT-test statistics conditional on $Q$ and $R$ under CSR independence. In this article, I use the empirically estimated expected values of $Q$ and $R$ under CSR independence pattern to remove the conditioning of NNCT-tests (such a correction is called the \emph{QR-adjustment}, henceforth). I present a Monte Carlo simulation study to compare the conditional NNCT-tests and QR-adjusted tests under CSR independence and segregation and association alternatives. I demonstrate that QR-adjustment does not significantly improve the empirical size estimates under CSR independence and power estimates under segregation or association alternatives. For illustrative purposes, I apply the conditional and empirically corrected tests on two example data sets.
研究动机与目标
- 为解决在CSR独立性下NNCT检验的条件性问题,即检验统计量依赖于随机的Q和R值。
- 开发一种实用的校正方法(QR校正),用CSR下经验估计的期望值替代观测到的Q和R值。
- 评估QR校正是否能提升NNCT检验在检测空间分离或关联方面的有效性与性能。
- 在CSR独立性及替代模式下,比较条件(未校正)与无条件(QR校正)NNCT检验。
提出的方法
- 通过大量蒙特卡洛模拟,经验估计CSR独立性下共享最近邻(Q)和反射最近邻(R)的期望值。
- 将NNCT检验统计量中的观测Q和R值替换为估计的期望值,以构建QR校正后的检验统计量。
- 将Dixon的整体检验和Ceyhan的三种新型分离检验的QR校正版本应用于NNCT。
- 通过蒙特卡洛模拟,比较在CSR独立性和分离/关联替代模式下,未校正与QR校正检验的经验大小和统计功效估计值。
- 采用相同的NNCT结构和检验统计量(如卡方检验),但将条件基于估计的E[Q]和E[R],而非观测到的Q和R。
- 将未校正和QR校正检验应用于真实和人工数据集,以说明其实际影响。
实验结果
研究问题
- RQ1QR校正是否能改善NNCT检验在CSR独立性原假设下的经验大小?
- RQ2QR校正是否能增强NNCT检验在分离或关联替代假设下的统计功效?
- RQ3在CSR及替代模式下,QR校正与未校正NNCT检验在第一类与第二类错误率方面有何差异?
- RQ4使用Q和R的经验估计期望值是否能有效替代观测值,而不导致推断偏差?
- RQ5在何种条件下,QR校正可能与未校正检验得出不同结论?
主要发现
- QR校正对NNCT检验在CSR独立性原假设下的经验大小无显著影响。
- 在分离或关联替代假设下,NNCT检验的统计功效在QR校正后基本保持不变。
- QR校正后检验统计量略有下降,但变化幅度不足以在所研究的示例中改变结论。
- 在含100个点(50个X,50个Y)的人工数据集中,未校正与QR校正检验均未能拒绝CSR独立性(p值 > 0.05)。
- 对于沼泽树数据,未校正与QR校正检验得出了相同结论:存在强烈的物种分离证据。
- 未校正与QR校正版本的p值差异极小,表明在典型情况下,QR校正不会显著改变推断结果。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。