[论文解读] Determining the Number of Factors in High-dimensional Generalised Latent Factor Models
本文提出了一种一致的信息准则,用于确定高维广义潜变量因子模型中的因子数量,该模型可处理混合数据类型和缺失值。该方法在高维渐近条件下确保了准确的因子选择,具有理论误差界,并在模拟数据和真实数据上表现出强劲的实证性能,包括对艾森克人格问卷中三因子结构的验证。
As a generalisation of the classical linear factor model, generalised latent factor models are a useful tool for analysing multivariate data of different types, including binary choices and counts. In this paper, we propose an information criterion to determine the number of factors in generalised latent factor models. The consistency of the proposed information criterion is established under a high-dimensional setting where both the sample size and the number of manifest variables grow to infinity and data may have many missing values. To establish this consistency result, an error bound is established for the parameter estimates that improves the existing results and may be of independent theoretical interest. Simulation shows that the proposed criterion has good finite sample performance. An application to Eysenck's personality questionnaire confirms the three-factor structure of this personality survey.
研究动机与目标
- 解决在具有混合类型多变量数据的高维广义潜变量因子模型中确定因子数量的挑战。
- 开发一种在样本量和可观测变量数量均趋于无穷大时仍保持一致的因子选择准则。
- 考虑高维数据中普遍存在大量缺失值的情况,这在实际应用中十分常见。
- 建立更紧致的参数估计误差界,从而在理论上增强高维因子建模的贡献。
- 在模拟数据和真实世界人格调查数据(艾森克问卷)上对方法进行实证验证。
提出的方法
- 基于似然近似提出一种信息准则,专用于具有非正态分布和混合类型数据(如二值变量、计数变量)的广义潜变量因子模型。
- 在高维渐近条件下推导该准则的一致性,其中样本量 n 和可观测变量数 p 均趋于无穷大。
- 建立新的参数估计误差界,优于现有结果,从而增强理论稳健性。
- 通过基于似然的估计方法处理缺失数据,使该方法可应用于不完整数据集。
- 采用惩罚似然方法平衡模型拟合度与复杂度,从而在渐近意义上偏好真实的因子数量。
实验结果
研究问题
- RQ1在具有混合数据类型的高维广义潜变量因子模型中,信息准则能否一致地选择出正确的因子数量?
- RQ2当数据包含大量缺失值时,所提出的准则在有限样本中的表现如何?
- RQ3参数估计误差界改进的程度在多大程度上提升了因子选择的可靠性?
- RQ4该方法能否正确识别艾森克人格问卷中既定的三因子结构?
主要发现
- 所提出的的信息准则在高维渐近条件下具有一致性,即使存在缺失数据,也能正确选择真实因子数量。
- 在模拟实验中,该方法在有限样本下表现出良好的性能,能准确恢复各种数据配置下的真实因子数量。
- 所推导的参数估计误差界比现有结果更紧致,表明理论稳定性得到提升。
- 对艾森克人格问卷的实际应用验证了三因子结构的存在,支持了该方法的有效性。
- 该方法能有效处理混合数据类型(如二值变量和计数变量)以及缺失值,且无需进行数据插补。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。