Skip to main content
QUICK REVIEW

[论文解读] Bayesian analysis for a class of beta mixed models

Wagner Hugo Bonat, Paulo Justiniano Ribeiro|arXiv (Cornell University)|Jan 13, 2014
Statistical Methods and Bayesian Inference参考文献 35被引用 9
一句话总结

本文提出了一种基于集成嵌套拉普拉斯近似(INLA)的贝叶斯推断方法,用于贝塔混合模型,展示了其在建模比例、比率等有界连续响应变量时的计算效率和准确性。结果表明,INLA在计算时间显著减少的同时,能够产生与MCMC和似然方法相当的结果,使其成为在层次化数据结构中进行广泛模型比较和敏感性分析的理想选择。

ABSTRACT

Generalized linear mixed models (GLMM) encompass large class of statistical models, with a vast range of applications areas. GLMM extends the linear mixed models allowing for different types of response variable. Three most common data types are continuous, counts and binary and standard distributions for these types of response variables are Gaussian, Poisson and Binomial, respectively. Despite that flexibility, there are situations where the response variable is continuous, but bounded, such as rates, percentages, indexes and proportions. In such situations the usual GLMM's are not adequate because bounds are ignored and the beta distribution can be used. Likelihood and Bayesian inference for beta mixed models are not straightforward demanding a computational overhead. Recently, a new algorithm for Bayesian inference called INLA (Integrated Nested Laplace Approximation) was proposed.INLA allows computation of many Bayesian GLMMs in a reasonable amount time allowing extensive comparison among models. We explore Bayesian inference for beta mixed models by INLA. We discuss the choice of prior distributions, sensitivity analysis and model selection measures through a real data set. The results obtained from INLA are compared with those obtained by an MCMC algorithm and likelihood analysis. We analyze data from an study on a life quality index of industry workers collected according to a hierarchical sampling scheme. Results show that the INLA approach is suitable and faster to fit the proposed beta mixed models producing results similar to alternative algorithms and with easier handling of modeling alternatives. Sensitivity analysis, measures of goodness of fit and model choice are discussed.

研究动机与目标

  • 开发并评估一种适用于比例、比率和指数等有界连续响应变量的贝塔混合模型贝叶斯框架。
  • 比较INLA与传统的MCMC和基于似然的推断方法在拟合贝塔混合模型时的性能。
  • 评估后验分布对分散参数(φ)和随机效应精度(τ)先验分布的敏感性。
  • 评估模型选择准则——LML、DIC和CPO——在真实世界层次化数据集中识别最佳拟合模型的能力。
  • 基于一项生活品质指数研究的真实数据,为贝塔混合模型中的先验设定和模型诊断提供实际指导。

提出的方法

  • 利用集成嵌套拉普拉斯近似(INLA)实现贝塔混合模型的快速贝叶斯推断,避免MCMC带来的计算负担。
  • 采用具有高斯随机效应的层次化建模结构,以反映数据中的聚类效应,例如员工嵌套于公司和州内。
  • 使用具有logit链接函数的线性预测器对贝塔分布的均值进行建模,并允许精度参数(φ)通过协变量变化。
  • 对分散参数(φ)和随机效应精度(τ)采用弱信息性的伽马先验,并通过Hellinger散度进行敏感性分析。
  • 使用包括对数边际似然(LML)、偏差信息准则(DIC)和条件预测有序统计量(CPO)在内的模型比较准则进行模型选择。
  • 通过与MCMC和基于似然的推断结果进行对比,验证了各方法间的一致性。

实验结果

研究问题

  • RQ1与MCMC和似然方法相比,INLA能否提供准确且计算高效的贝叶斯推断?
  • RQ2后验估计对分散参数(φ)和随机效应精度(τ)的先验选择有多敏感?
  • RQ3在层次化数据结构中,LML、DIC或CPO中哪一个准则最能识别最优的贝塔混合模型?
  • RQ4对φ和τ采用不同的先验分布在多大程度上影响后验分布和模型结论?
  • RQ5在层次化工业工人数据集中,公司规模和平均收入等协变量如何影响生活品质指数?

主要发现

  • INLA生成的贝塔混合模型后验估计与MCMC和似然推断结果高度一致,验证了其准确性。
  • φ的先验与后验分布之间的Hellinger距离从0.6降至0.1628,表明对先验选择的敏感性较低;τ的该距离从0.6降至0.2827,表明具有中等敏感性。
  • 尽管在极端先验下τ的后验均值变化达52.05%,但回归系数在各模型间变化均小于0.5%,保持稳定。
  • 通过LML、DIC和CPO选出的最终模型一致,表明模型选择具有稳健性。
  • 公司规模和平均收入被发现是生活品质指数的显著预测因子,随机截距项有效捕捉了州层面的变异。
  • INLA的计算时间远低于MCMC,使其能够高效探索多种模型设定。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。