[论文解读] Composite empirical likelihood for multisample clustered data
本文提出了一种基于复合经验似然的多组聚类数据监控检验方法,采用非参数随机效应模型处理组内相关性,并利用密度比模型整合多组信息。该方法采用基于聚类的自助法,在存在相关性的情况下仍能保持分位数监控的精确第一类错误控制,并实现对木材强度分布低百分位数的渐近有效推断。
In many applications, data cluster. Failing to take the cluster structure into consideration generally leads to underestimated variances of point estimators and inflated type I errors in hypothesis tests. Many circumstance-dependent approaches have been developed to handle clustered data. A working covariance matrix may be used in generalized estimating equations. One may throw out the cluster structure and use only the cluster means, or explicitly model the cluster structure. Our interest is the case where multiple samples of clustered data are collected, and the population quantiles are particularly important. We develop a composite empirical likelihood for clustered data under a density ratio model. This approach avoids parametric assumptions on the population distributions or the cluster structure. It efficiently utilizes the common features of the multiple populations and the exchangeability of the cluster members. We also develop a cluster-based bootstrap method to provide valid variance estimation and to control the type I errors. We examine the performance of the proposed method through simulation experiments and illustrate its usage via a real-world example.
研究动机与目标
- 解决现有非参数检验在数据聚类时检测木材强度低百分位数变化的局限性。
- 开发一种在组内相关性下控制第一类错误的稳健监控检验方法,这是林业和环境数据中的常见问题。
- 将复合经验似然与非参数随机效应建模相结合,避免对复杂相关结构做显式假设。
- 利用密度比模型整合多组信息,以提高效率和推断准确性。
- 提出一种基于聚类的自助程序,保持聚类结构,确保分位数变化检测的有效推断。
提出的方法
- 使用非参数随机效应模型以考虑多组数据中的组内相关性。
- 应用复合经验似然以避免对复杂相关结构进行显式建模。
- 采用密度比模型整合多组信息,提高估计效率。
- 开发一种基于聚类的自助程序,保持聚类结构,以估计分位数估计量的抽样分布。
- 推导在自助法下分位数估计量的Bahadur表示,支持渐近正态性与推断。
- 利用自助法构造分位数的置信区间并进行假设检验,理论依据来自渐近结果。
实验结果
研究问题
- RQ1如何使多组聚类数据中木材强度低百分位数变化的监控检验对组内相关性具有鲁棒性?
- RQ2复合经验似然结合基于聚类的自助法是否能在相关性下保持准确的第一类错误率?
- RQ3当数据聚类时,与现有非参数检验(如Wilcoxon检验、Kolmogorov检验)相比,所提方法在统计功效和错误控制方面表现如何?
- RQ4在所提出的复合似然与自助法框架下,分位数估计量的渐近行为如何?
- RQ5密度比模型在整合多组信息时,能在多大程度上提高估计效率?
主要发现
- 所提出的基于聚类的自助程序在组内相关性下成功控制了分位数监控的第一类错误率,而传统非参数检验则不能。
- 分位数的复合经验似然估计量具有高效率,并具有Bahadur表示,误差率为 $ O_p(n^{-3/4} \\_log^{3/4}n) $。
- 在自助法下,分位数估计量的联合渐近正态性得以建立,支持对多组数据的有效推断。
- 即使聚类数量适中,该方法在模拟研究中也表现出良好性能。
- 分位数的自助置信区间在聚类相关性下具有良好的校准性与鲁棒性。
- 对木材强度监控的真实数据应用证实了该方法的实际效用,并在聚类存在时优于现有检验方法。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。