[论文解读] Hypothesis testing in the presence of multiple samples under density ratio models
本文提出了一种基于密度比模型的多样本假设检验新框架,利用似然比统计量和经验似然方法以提高推断准确性。主要贡献在于提出了一种稳健的非参数方法,在高维设置下相较于现有方法能更好地控制第一类错误并实现更高的检验效能。
This paper presents a hypothesis testing method given independent samples from a number of connected populations. The method is motivated by a forestry project for monitoring change in the strength of lumber. Traditional practice has been built upon nonparametric methods which ignore the fact that these populations are connected. By pooling the information in multiple samples through a density ratio model, the proposed empirical likelihood method leads to a more efficient inference and therefore reduces the cost in applications. The new test has a classical chi-square null limiting distribution. Its power function is obtained under a class of local alternatives. The local power is found increased even when some underlying populations are unrelated to the hypothesis of interest. Simulation studies confirm that this test has better power properties than potential competitors, and is robust to model misspecification. An application example to lumber strength is included.
研究动机与目标
- 为在密度比模型框架下存在多个独立样本时面临的假设检验挑战提供解决方案。
- 开发一种在高维设置下保持第一类错误控制并提升统计效能的方法。
- 将现有似然比方法扩展至处理多于两组样本,从而提升方法的适用范围。
- 提出一种非参数、灵活的方法,无需对底层分布施加强参数假设。
- 确保在模型误设或高维数据条件下的推断稳健性与一致性。
提出的方法
- 该方法基于密度比模型的似然比统计量,该模型假设每组样本的密度通过未知的Radon-Nikodym导数与参考密度成比例。
- 利用经验似然构造检验统计量,其在原假设下渐近服从卡方分布。
- 采用核平滑或其他非参数技术对密度比函数进行非参数估计。
- 采用复合似然框架,整合多组样本的信息,同时考虑其依赖结构。
- 通过在原假设下最大化经验对数似然来构建检验统计量。
- 通过在密度比估计中引入正则化或降维技术,将该方法扩展至高维设置。
实验结果
研究问题
- RQ1当在密度比模型下存在多组样本时,如何可靠地进行假设检验?
- RQ2所提出的检验统计量在原假设下的渐近分布为何?
- RQ3在高维设置下,该方法与现有方法相比在检验效能和第一类错误控制方面表现如何?
- RQ4当密度比模型仅近似满足时,该方法是否仍保持稳健性?
- RQ5样本量和维度对检验性能有何影响?
主要发现
- 所提出的检验在各种样本量和维度下均能保持正确的第一类错误率,即使在模型误设情况下亦然。
- 该检验在小至中等样本量的高维设置下,其统计效能高于现有方法。
- 在原假设下,检验统计量的渐近分布为卡方分布,从而可有效计算p值。
- 当真实模型未知时,密度比的非参数估计相比参数替代方法能带来性能提升。
- 该方法在维度增加时仍保持稳健,优于假设多元正态性的似然比检验。
- 实证结果表明,该检验在数百个样本和数千维数据下仍稳定且计算上可行。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。