[论文解读] Data Integration with High Dimensionality
本文提出了一种用于高维数据整合的伪似然信息准则,通过结合多个实验的边际似然,当真实预测变量数量随样本量增长时,选择具有信息量的预测变量对象。在常规条件下,该方法建立了选择一致性,通过二次型的大偏差界,将贝叶斯信息准则推广至无界模型规模,并提供理论保证。
We consider a problem of data integration. Consider determining which genes affect a disease. The genes, which we call predictor objects, can be measured in different experiments on the same individual. We address the question of finding which genes are predictors of disease by any of the experiments. Our formulation is more general. In a given data set, there are a fixed number of responses for each individual, which may include a mix of discrete, binary and continuous variables. There is also a class of predictor objects, which may differ within a subject depending on how the predictor object is measured, i.e., depend on the experiment. The goal is to select which predictor objects affect any of the responses, where the number of such informative predictor objects or features tends to infinity as sample size increases. There are marginal likelihoods for each way the predictor object is measured, i.e., for each experiment. We specify a pseudolikelihood combining the marginal likelihoods, and propose a pseudolikelihood information criterion. Under regularity conditions, we establish selection consistency for the pseudolikelihood information criterion with unbounded true model size, which includes a Bayesian information criterion with appropriate penalty term as a special case. Simulations indicate that data integration improves upon, sometimes dramatically, using only one of the data sources.
研究动机与目标
- 解决预测变量对象(例如基因)通过多种实验测量、响应类型混合(连续型、二值型、离散型)的数据整合问题。
- 开发一种在真实预测变量数量随样本量趋于无穷时仍保持一致的模型选择准则。
- 将不同实验的边际似然整合到统一的伪似然框架中,以实现联合推断。
- 在一般正则条件下(包括模型误设情况)建立所提信息准则的理论一致性。
- 将现有信息准则(如BIC)扩展至模型规模发散且预测变量对象高维化的场景。
提出的方法
- 通过结合K个不同实验的边际似然构建伪似然,每个实验以不同方法测量同一组预测变量对象。
- 推导出一种带有惩罚项的伪似然信息准则,旨在处理无界的真实模型规模。
- 应用大偏差理论处理二次型,推导出尾部概率的紧致上界,替代不可靠的渐近卡方近似。
- 利用矩生成函数技术和累积量有界性,控制伪似然得分及其导数的偏离。
- 采用邦弗朗尼不等式和集中不等式,控制模型空间中所有模型的最大偏离。
- 通过证明错误模型被选中的概率随样本量增加而趋于零,建立一致性,即使真实预测变量数量随n增长亦成立。
实验结果
研究问题
- RQ1当真实预测变量数量随样本量增长时,基于伪似然的信息准则是否能实现模型选择一致性?
- RQ2在模型规模发散且高维化的设定下,信息准则中的惩罚项应如何设计以保持一致性?
- RQ3当渐近卡方近似失效时,需要哪些理论工具来界定伪似然比统计量的尾部概率?
- RQ4与单一数据源相比,跨多个实验的数据整合在多大程度上能提升模型选择性能?
- RQ5在何种正则条件下,所提准则能在模型误设情况下一致地选择真实模型?
主要发现
- 所提伪似然信息准则在常规条件下,即使真实预测变量数量随样本量增长,仍能实现选择一致性。
- 当惩罚项选择为 $ \gamma_n = 6w(1+\gamma)\log(p_n) $ 或 $ \gamma_n = 6w(\log p_n + \log \log p_n) $ 时,该准则被证明具有一致性,确保模型选择概率收敛至1。
- 对二次型的大偏差界提供了紧致的有限样本尾部概率上界,替代了不可靠的渐近卡方近似。
- 模拟结果显示,该方法在性能上优于单一数据源分析,数据整合显著提升了模型选择的准确性。
- 理论结果将贝叶斯信息准则推广至无界模型规模的高维设定,推广了以往局限于有界真实模型规模的研究。
- 一致性证明依赖于利用集中不等式和矩界控制所有模型中伪似然得分及其导数的最大偏离。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。