[论文解读] Conditional Inference for Multivariate Generalised Linear Mixed Models
本文提出了一种用于多元广义线性混合模型(MGLMMs)的新型条件推断方法,通过预测随机效应避免了对似然函数的积分。该方法将GLMMs扩展至非正态随机效应(例如,多元t分布)和散度模型,实现了对具有复杂依赖结构的多样化响应类型的灵活建模,同时保持了渐近效率和计算可行性。
We propose a method for inference in generalised linear mixed models (GLMMs) and several extensions of these models. First, we extend the GLMM by allowing the distribution of the random components to be non-Gaussian, that is, assuming an absolutely continuous distribution with respect to the Lebesgue measure that is symmetric around zero, unimodal and with finite moments up to fourth-order. Second, we allow the conditional distribution to follow a dispersion model instead of exponential dispersion models. Finally, we extend these models to a multivariate framework where multiple responses are combined by imposing a multivariate absolute continuous distribution on the random components representing common clusters of observations in all the marginal models. Maximum likelihood inference in these models involves evaluating an integral that often cannot be computed in closed form. We suggest an inference method that predicts values of random components and does not involve the integration of conditional likelihood quantities. The multivariate GLMMs that we studied can be constructed with marginal GLMMs of different statistical nature, and at the same time, represent complex dependence structure providing a rather flexible tool for applications.
研究动机与目标
- 开发一种无需对条件似然函数进行高维积分的GLMMs无似然推断方法。
- 通过允许具有有限四阶矩的非高斯、对称、单峰随机效应,扩展标准GLMMs。
- 将条件分布族的范围从指数分散族推广至更广泛的散度模型。
- 构建多元GLMMs,使得边际模型可以属于不同的统计分布类型(例如,二项分布、泊松分布、伽马分布),同时共享一个共同的多元随机效应结构。
- 提供一种计算高效且渐近有效的推断框架,适用于具有异质响应的复杂真实世界多元数据。
提出的方法
- 基于条件估计方程的推断函数,绕过了对随机效应积分的需求。
- 通过条件估计方法直接预测随机效应值,避免了对似然函数的数值积分。
- 采用多元连续分布(例如,多元正态分布或t分布)作为随机效应分布,允许在不同聚类间实现灵活的依赖结构。
- 将模型框架扩展至包含散度模型,而非仅限于指数分散族,从而提升模型灵活性。
- 在附录A.4中应用了拉普拉斯近似的多元扩展,用于与所提方法进行比较和验证。
- 使用Hermite近似和基于模拟的评估方法,评估不同随机效应分布和聚类规模下的性能表现。
实验结果
研究问题
- RQ1能否为GLMMs开发一种条件推断方法,避免对随机效应进行积分,同时保持渐近效率?
- RQ2非高斯随机效应(例如,多元t分布)如何影响多元GLMMs中推断的性能与稳健性?
- RQ3在MGLMMs中,边际模型在分布族和链接函数方面可以有多大差异,同时仍能维持一致的多元依赖结构?
- RQ4与Breslow和Clayton(1993)的拉普拉斯近似等现有方法相比,所提推断方法在估计量的偏差、标准误和正态性方面表现如何?
- RQ5该方法能否扩展至指数族以外的散度模型?这对模型灵活性和推断准确性有何影响?
主要发现
- 当随机效应为高斯分布时,所提条件推断方法的性能与Breslow和Clayton(1993)的拉普拉斯近似相当,偏差和标准误相近。
- 即使随机效应服从重尾分布(如多元t分布),该方法在有限样本下仍表现出良好的性质。
- 模拟研究显示,参数估计量的抽样分布近似正态,Shapiro-Wilk检验的p值范围为0.15至0.85,表明正态近似良好。
- 在不同聚类规模(q = 10, 50, 100)下,参数估计偏差均保持较低水平,大多数参数的绝对偏差值低于0.05。
- 该方法成功处理了具有混合边际分布族(如二项分布、泊松分布、伽马分布)的多元GLMMs,实现了对异质响应的联合建模。
- 附录A.4中多元拉普拉斯近似的扩展结果证实,在标准正则性条件下,所提推断框架具有一致性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。