Skip to main content
QUICK REVIEW

[Paper Review] Data Integration with High Dimensionality

Xin Gao, Raymond J. Carroll|arXiv (Cornell University)|Oct 3, 2016
Gene expression and cancer classification3 citations
TL;DR

This paper proposes a pseudolikelihood information criterion for high-dimensional data integration, combining marginal likelihoods from multiple experiments to select informative predictor objects when the number of true predictors grows with sample size. It establishes selection consistency under regularity conditions, extending Bayesian information criteria to unbounded model sizes with theoretical guarantees via large deviation bounds on quadratic forms.

ABSTRACT

We consider a problem of data integration. Consider determining which genes affect a disease. The genes, which we call predictor objects, can be measured in different experiments on the same individual. We address the question of finding which genes are predictors of disease by any of the experiments. Our formulation is more general. In a given data set, there are a fixed number of responses for each individual, which may include a mix of discrete, binary and continuous variables. There is also a class of predictor objects, which may differ within a subject depending on how the predictor object is measured, i.e., depend on the experiment. The goal is to select which predictor objects affect any of the responses, where the number of such informative predictor objects or features tends to infinity as sample size increases. There are marginal likelihoods for each way the predictor object is measured, i.e., for each experiment. We specify a pseudolikelihood combining the marginal likelihoods, and propose a pseudolikelihood information criterion. Under regularity conditions, we establish selection consistency for the pseudolikelihood information criterion with unbounded true model size, which includes a Bayesian information criterion with appropriate penalty term as a special case. Simulations indicate that data integration improves upon, sometimes dramatically, using only one of the data sources.

Motivation & Objective

  • To address data integration problems where predictor objects (e.g., genes) are measured via multiple experiments with mixed response types (continuous, binary, discrete).
  • To develop a model selection criterion that remains consistent when the number of true predictors tends to infinity with sample size.
  • To combine marginal likelihoods from different experiments into a unified pseudolikelihood framework for joint inference.
  • To establish theoretical consistency of the proposed information criterion under general regularity conditions, including model misspecification.
  • To extend existing information criteria—such as BIC—to settings with diverging model size and high-dimensional predictor objects.

Proposed method

  • Formulates a pseudolikelihood by combining marginal likelihoods from K different experiments, each measuring the same set of predictor objects via distinct methods.
  • Derives a pseudolikelihood information criterion with a penalty term designed to handle unbounded true model size.
  • Applies large deviation theory to quadratic forms to derive sharp upper bounds on tail probabilities, replacing asymptotic chi-square approximations.
  • Uses cumulant boundedness and moment generating function techniques to control the deviation of pseudo-likelihood scores and their derivatives.
  • Employs Bonferroni inequalities and concentration inequalities to bound the maximum deviation across all models in the model space.
  • Establishes consistency by showing that the probability of selecting a wrong model vanishes as sample size increases, even when the number of true predictors grows with n.

Experimental results

Research questions

  • RQ1Can a pseudolikelihood-based information criterion achieve model selection consistency when the number of true predictors increases with sample size?
  • RQ2How should the penalty term in an information criterion be designed to maintain consistency in high-dimensional settings with diverging model size?
  • RQ3What theoretical tools are required to bound the tail probabilities of pseudolikelihood ratio statistics when asymptotic chi-square approximations fail?
  • RQ4To what extent does data integration across multiple experiments improve model selection performance compared to using a single data source?
  • RQ5Under what regularity conditions does the proposed criterion consistently select the true model in the presence of model misspecification?

Key findings

  • The proposed pseudolikelihood information criterion achieves selection consistency even when the number of true predictors grows with sample size, under regularity conditions.
  • The criterion is shown to be consistent when the penalty term is chosen as $ \gamma_n = 6w(1+\gamma)\log(p_n) $ or $ \gamma_n = 6w(\log p_n + \log \log p_n) $, ensuring the model selection probability converges to one.
  • Large deviation bounds on quadratic forms provide sharp, finite-sample upper bounds on tail probabilities, replacing unreliable asymptotic chi-square approximations.
  • The method outperforms single-source analysis in simulations, with data integration leading to dramatic improvements in model selection accuracy.
  • Theoretical results extend the Bayesian information criterion to high-dimensional settings with unbounded model size, generalizing prior work limited to bounded true model size.
  • The consistency proof relies on controlling the maximum deviation of pseudo-likelihood scores and their derivatives across all models using concentration inequalities and moment bounds.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.