[论文解读] Weighted scores method for longitudinal ordinal data
本文提出加权得分法作为广义估计方程(GEE)在纵向有序数据中的稳健、高效替代方法,避免将有序响应转换为二值指标。通过利用离散化的多元正态工作模型,该方法降低了计算负担,避免了复杂的相关矩阵,从而在具有大量类别时实现更快、更稳定的估计。
Extending generalized estimating equations (GEE) to ordinal response data requires a conversion of the ordinal response to a vector of binary category indicators. That leads to a rather complicated association structure, and the introduction of large matrices when the number of categories and dimension of the cluster are large. To allow a richer specification of working correlation assumptions, we adopt the weighted scores method which is essentially an extension of the GEE approach, since it can also be applied to families that are not in the GLM class. The weighted scores method stems from the lack of a theoretically sound methodology for analyzing multivariate discrete data based only on moments up to second order and it is robust to dependence and nearly as efficient as maximum likelihood. There is no need to convert the ordinal response to binary indicators, thus the weight matrices have smaller dimensions and it is not necessary to guess the correlations of indicator variables for different categories. We focus on important issues that would interest the data analyst, such as choice of the structure of the correlation matrix and of explanatory variables, comparison of results obtained from our methods versus GEE, and insights provided by our method that would be missed with the GEE method. Our modelling framework is implemented in the package weightedScores within the open source statistical environment R.
研究动机与目标
- 解决将GEE扩展至具有大量类别的有序响应时计算效率低下和复杂度高的问题。
- 消除将有序结果转换为二值指标的需要,因为这会显著增加工作相关矩阵的维度。
- 提供一种基于二阶矩的稳健方法,其效率接近最大似然法,但相较于GEE在K或d较大时更具可扩展性。
- 在具有高维聚类或大量类别的纵向有序数据中实现可靠推断,例如临床评分系统。
- 为现有GEE方法提供一种实用、计算可行的替代方案,这些方法常因收敛问题和缓慢的矩阵运算而受限。
提出的方法
- 该方法使用离散化的多元正态分布作为工作模型,以定义单变量得分函数的权重。
- 通过基于边际分布和成对关联的权重函数构造估计方程,避免将响应转换为二值指标。
- 权重矩阵基于潜在变量框架下二元有序响应的二阶矩推导得出。
- 通过求解加权估计方程来估计回归参数和关联参数,其方差-协方差矩阵采用经验异方差-异质性稳健估计器计算。
- 支持灵活的相关结构选择,并能高效处理大量类别(K)和聚类规模(d)。
- 该方法已通过R包'weightedScores'实现,便于在生物统计学研究中实际应用。
实验结果
研究问题
- RQ1如何在不将有序响应转换为二值指标的前提下,高效建模具有大量类别的纵向有序数据?
- RQ2与传统GEE相比,加权得分法在计算和统计方面具有哪些优势?
- RQ3与现有GEE方法相比,该方法在估计效率和收敛性方面表现如何?
- RQ4在高维设置下,加权得分法是否能揭示GEE所遗漏的见解?
- RQ5不同相关结构假设对纵向有序模型中参数估计和推断的影响是什么?
主要发现
- 加权得分法避免了将有序响应转换为K−1个二值指标,显著降低了工作相关矩阵的维度。
- 该方法在类别数K或聚类规模d较大时,表现出更高的计算速度和稳定性。
- 在二阶矩假设下,即使真实分布被错误设定,该方法仍能提供一致且近乎高效的估计量。
- 该方法对依赖结构具有鲁棒性,并在各种相关结构(包括交换相关和自回归相关)下保持良好性能。
- 实证比较表明,加权得分法产生的估计值与GEE相似,但收敛速度更快,计算负担更轻。
- R包'weightedScores'支持实际应用,具备变量选择、基于AIC/BIC的模型比较以及相关结构选择功能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。