[论文解读] How robust are Structural Equation Models to model miss-specification? A simulation study
本模拟研究利用 R 语言中的 lavaan 和 piecewiseSEM 评估结构方程模型(SEMs)在模型误设情况下的稳健性。研究发现,单一拟合指数无法可靠检测误设;相反,结合多种指标——尤其是 BIC——可提升模型选择效果,而充足的样本量(>100)对于检测效应及避免过拟合或欠拟合至关重要。
Structural Equation Models (SEMs) are routinely used in the analysis of empirical data by researchers from different scientific fields such as psychologists or economists. In some fields, such as in ecology, SEMs have only started recently to attract attention and thanks to dedicated software packages the use of SEMs has steadily increased. Yet, common analysis practices in such fields that might be transposed from other statistical techniques such as model acceptance or rejection based on p-value screening might be poorly fitted for SEMs especially when these models are used to confirm or reject hypotheses. In this simulation study, SEMs were fitted via two commonly used R packages: lavaan and piecewiseSEM. Five different data-generation scenarios were explored: (i) random, (ii) exact, (iii) shuffled, (iv) underspecified and (v) overspecified. In addition, sample size and model complexity were also varied to explore their impact on various global and local model fitness indices. The results showed that not one single model index should be used to decide on model fitness but rather a combination of different model fitness indices is needed. The global chi-square test for lavaan or the Fisher's C statistic for piecewiseSEM were, in isolation, poor indicators of model fitness. In addition, the simulations showed that to achieve sufficient power to detect individual effects, adequate sample sizes are required. Finally, BIC showed good capacity to select models closer to the truth especially for more complex models. I provide, based on these results, a flowchart indicating how information from different metrics may be combined to reveal model strength and weaknesses. Researchers in scientific fields with little experience in SEMs, such as in ecology, should consider and accept these limitations.
研究动机与目标
- 评估常用 SEM 拟合指数在真实生态数据情景下检测模型误设的能力。
- 研究样本量和模型复杂度对全局与局部拟合指数性能的影响。
- 比较信息准则(AIC、BIC、HBIC)在选择更接近真实数据生成过程的模型方面的有效性。
- 为生态学及其他 SEM 经验有限的研究人员提供解释和改进模型拟合的实用指导。
- 强调依赖单一拟合统计量(如卡方检验或 Fisher’s C)进行 SEM 评估的局限性。
提出的方法
- 模拟五种数据生成情景:随机、精确、打乱、欠定和过定模型。
- 通过改变样本量和模型复杂度,评估其对 lavaan 和 piecewiseSEM 包中拟合指数的影响。
- 评估全局拟合指数(lavaan 中的卡方检验,piecewiseSEM 中的 Fisher’s C)和局部指数(路径 p 值、R 平方、条件独立性检验)。
- 使用信息准则(AIC、AICc、BIC、HBIC)比较不同情景下的模型选择表现。
- 在多元正态分布下生成数据,不包含层次结构或潜变量,以确保可比性。
- 开发了一个决策流程图,整合多种指标,以指导模型评估与改进。
实验结果
研究问题
- RQ1在不同数据生成情景下,全局拟合指数(如卡方检验、Fisher’s C)在检测模型误设方面的表现如何?
- RQ2样本量和模型复杂度在多大程度上影响 SEM 中检测单个路径效应的能力?
- RQ3在各种误设条件下,哪种信息准则(AIC、AICc、BIC、HBIC)最能识别出真实模型?
- RQ4局部拟合指数(如路径 p 值、R 平方、条件独立性检验)能否帮助诊断过拟合或欠拟合?
- RQ5哪种拟合指标组合能为 SEM 中的模型改进提供最可靠的指导?
主要发现
- 单一全局拟合指数(如 lavaan 中的卡方检验或 piecewiseSEM 中的 Fisher’s C)在孤立使用时无法可靠检测模型误设。
- BIC 在选择更接近真实数据生成过程的模型方面表现出最强能力,尤其在复杂模型和大样本量下。
- 样本量低于 100 时,检测单个路径效应的能力有限,凸显了在 SEM 应用中采用充足样本量的重要性。
- AICc 在模型选择中的表现随样本量呈现倒 U 型变化,而 BIC 和 HBIC 随样本量增大表现出渐近改善。
- 全局 p 值较高但显著路径较少且 R 平方较低的模型表明信号提取不足,提示可能存在过拟合或缺乏有意义的关系。
- 相反,全局拟合差但显著路径多且 R 平方高的模型提示可能存在欠拟合或方向假设错误,表明存在缺失关系或路径误设。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。