[论文解读] Variance partitioning in multilevel models for count data
本文为具有计数数据的多水平模型推导了精确的方差分解系数(VPC)和组内相关系数(ICC),将先前针对泊松模型的研究扩展至更灵活的负二项分布、三级模型及随机系数模型。研究提供了将总方差分解为组间与组内成分的代数表达式,使研究者能够量化过度离势计数结果(如学生缺勤)中的聚类效应。
A first step when fitting multilevel models to continuous responses is to explore the degree of clustering in the data. Researchers fit variance-component models and then report the proportion of variation in the response that is due to systematic differences between clusters. Equally they report the response correlation between units within a cluster. These statistics are popularly referred to as variance partition coefficients (VPCs) and intraclass correlation coefficients (ICCs). When fitting multilevel models to categorical (binary, ordinal, or nominal) and count responses, these statistics prove more challenging to calculate. For categorical response models, researchers appeal to their latent response formulations and report VPCs/ICCs in terms of latent continuous responses envisaged to underly the observed categorical responses. For standard count response models, however, there are no corresponding latent response formulations. More generally, there is a paucity of guidance on how to partition the variation. As a result, applied researchers are likely to avoid or inadequately report and discuss the substantive importance of clustering and cluster effects in their studies. A recent article drew attention to a little-known exact algebraic expression for the VPC/ICC for the special case of the two-level random-intercept Poisson model. In this article, we make a substantial new contribution. First, we derive exact VPC/ICC expressions for more flexible negative binomial models that allows for overdispersion, a phenomenon which often occurs in practice. Then we derive exact VPC/ICC expressions for three-level and random-coefficient extensions to these models. We illustrate our work with an application to student absenteeism.
研究动机与目标
- 为解决多水平计数数据模型中缺乏可靠的方差分解方法,特别是存在过度离势时的问题。
- 将先前仅限于线性与二值模型的方差分解技术扩展至计数响应变量领域,且不依赖潜在变量假设。
- 推导更复杂多水平结构(包括计数数据的三级模型与随机系数模型)中VPC与ICC的精确代数表达式。
- 为应用研究者提供实用工具,以评估聚类在计数数据研究中的实际重要性。
- 通过学生缺勤的实际案例说明该方法,展示其在公共卫生与社会科学研究中的实用性。
提出的方法
- 推导了考虑过度离势的两级随机截距负二项分布模型的精确方差分解系数(VPC)与组内相关系数(ICC)。
- 将推导扩展至计数数据的三级多水平模型,实现跨多个层级的方差分解。
- 推导了负二项分布模型的随机系数(随机斜率)扩展的VPC与ICC表达式,允许存在聚 cluster 特异性效应。
- 基于负二项分布的边际方差结构,使用精确代数表达式计算归属于高层级聚 cluster 的总方差比例。
- 将该方法应用于真实的学生缺勤数据集,演示VPC与ICC的计算与解释过程。
- 通过比较不同模型类型的结果验证该方法,显示其与理论预期的一致性。
实验结果
研究问题
- RQ1当存在过度离势时,如何在多水平计数数据模型中实现方差分解?
- RQ2在具有计数结果的三级多水平模型中,VPC与ICC的精确代数表达式是什么?
- RQ3如何利用负二项分布模型的随机系数扩展来实现计数数据的方差分解?
- RQ4聚类对计数数据总方差的影响是什么?如何实现有意义的量化?
- RQ5这些方差分解技术如何在现实世界的公共卫生与社会科学研究中应用与解释?
主要发现
- 推导出两级随机截距负二项分布模型的精确VPC与ICC表达式,使研究者能够精确量化过度离势计数数据中的聚类效应。
- 该方法成功扩展至三级多水平模型,实现了计数数据在多个分层结构中的方差分解。
- 推导出随机系数负二项分布模型的VPC与ICC表达式,可容纳聚 cluster 特异性斜率,提升模型灵活性。
- 在学生缺勤数据上的应用表明,缺勤率总方差的相当大比例可归因于校级聚类,凸显了情境效应的重要性。
- 所推导的表达式在代数上是精确的,不依赖潜在变量近似,因此比以往的临时方法更具可靠性。
- 结果表明,忽略过度离势或使用错误的方差分解方法,可能导致对计数数据中聚类效应的误解。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。