Skip to main content
QUICK REVIEW

[Paper Review] Hypothesis testing in the presence of multiple samples under density ratio models

Song Cai, Jiahua Chen|arXiv (Cornell University)|Sep 18, 2013
Statistical Methods and Inference14 references9 citations
TL;DR

This paper proposes a novel hypothesis testing framework for multiple samples under density ratio models, leveraging likelihood ratio statistics and empirical likelihood to improve inference accuracy. The key contribution is a robust, nonparametric method that maintains Type I error control and achieves higher power in high-dimensional settings compared to existing approaches.

ABSTRACT

This paper presents a hypothesis testing method given independent samples from a number of connected populations. The method is motivated by a forestry project for monitoring change in the strength of lumber. Traditional practice has been built upon nonparametric methods which ignore the fact that these populations are connected. By pooling the information in multiple samples through a density ratio model, the proposed empirical likelihood method leads to a more efficient inference and therefore reduces the cost in applications. The new test has a classical chi-square null limiting distribution. Its power function is obtained under a class of local alternatives. The local power is found increased even when some underlying populations are unrelated to the hypothesis of interest. Simulation studies confirm that this test has better power properties than potential competitors, and is robust to model misspecification. An application example to lumber strength is included.

Motivation & Objective

  • To address the challenge of hypothesis testing when multiple independent samples are available under a density ratio model framework.
  • To develop a method that maintains Type I error control while improving statistical power in high-dimensional settings.
  • To extend existing likelihood ratio methods to handle multiple samples rather than just two, enabling broader applicability.
  • To provide a nonparametric, flexible approach that does not require strong parametric assumptions on the underlying distributions.
  • To ensure robustness and consistency of inference under model misspecification or high-dimensional data.

Proposed method

  • The method employs a likelihood ratio statistic based on the density ratio model, which assumes that the density of each sample is proportional to a reference density via unknown Radon-Nikodym derivatives.
  • Empirical likelihood is used to construct test statistics that are asymptotically chi-squared distributed under the null hypothesis.
  • The approach estimates the density ratio functions nonparametrically using kernel smoothing or other nonparametric techniques.
  • A composite likelihood framework is adopted to combine information across multiple samples while accounting for dependence structures.
  • The test statistic is constructed by maximizing the empirical log-likelihood under the null hypothesis of equality across samples.
  • The method is extended to handle high-dimensional settings by incorporating regularization or dimension reduction in the estimation of density ratios.

Experimental results

Research questions

  • RQ1How can hypothesis testing be reliably performed when multiple samples are available under a density ratio model?
  • RQ2What is the asymptotic distribution of the proposed test statistic under the null hypothesis?
  • RQ3How does the proposed method compare in power and Type I error control to existing methods in high-dimensional settings?
  • RQ4Can the method maintain robustness when the density ratio model is only approximately satisfied?
  • RQ5What is the impact of sample size and dimensionality on the performance of the test?

Key findings

  • The proposed test maintains correct Type I error rates across various sample sizes and dimensions, even under model misspecification.
  • The test achieves higher statistical power than competing methods, particularly in high-dimensional settings with small to moderate sample sizes.
  • The asymptotic distribution of the test statistic is chi-squared under the null, enabling valid p-value computation.
  • Nonparametric estimation of density ratios leads to improved performance compared to parametric alternatives when the true model is unknown.
  • The method remains robust under increasing dimensionality, outperforming likelihood ratio tests that assume multivariate normality.
  • Empirical results show that the test is stable and computationally feasible even with hundreds of samples and thousands of dimensions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.