[论文解读] High Dimensional Latent Panel Quantile Regression with an Application to Asset Pricing
本文提出了一种高维潜在面板分位数回归模型,通过使用ADMM结合ℓ₁范数与核范数正则化,联合估计资产定价中的稀疏特征与低秩潜在因子。该方法在时间依赖性和重尾误差下,一致地估计了稀疏系数与低秩因子结构,揭示了不同收益分位数下的风险暴露差异。
We propose a generalization of the linear panel quantile regression model to accommodate both \ extit{sparse} and \ extit{dense} parts: sparse means while the number of covariates available is large, potentially only a much smaller number of them have a nonzero impact on each conditional quantile of the response variable; while the dense part is represent by a low-rank matrix that can be approximated by latent factors and their loadings. Such a structure poses problems for traditional sparse estimators, such as the $\\ell_1$-penalised Quantile Regression, and for traditional latent factor estimator, such as PCA. We propose a new estimation procedure, based on the ADMM algorithm, consists of combining the quantile loss function with $\\ell_1$ \ extit{and} nuclear norm regularization. We show, under general conditions, that our estimator can consistently estimate both the nonzero coefficients of the covariates and the latent low-rank matrix. Our proposed model has a "Characteristics + Latent Factors" Asset Pricing Model interpretation: we apply our model and estimator with a large-dimensional panel of financial data and find that (i) characteristics have sparser predictive power once latent factors were controlled (ii) the factors and coefficients at upper and lower quantiles are different from the median.
研究动机与目标
- 解决在高维面板数据中联合建模稀疏公司特征与密集潜在因子的资产定价挑战。
- 克服传统ℓ₁惩罚分位数回归与主成分分析(PCA)在稀疏与密集分量共存时的局限性。
- 在时间依赖性、重尾分布及未知潜在因子结构下,开发一致估计量。
- 提供一个统一框架,以捕捉资产收益不同分位数下的异质性风险暴露。
- 通过区分基于特征的风险驱动因素与基于因子的风险驱动因素,提升经济可解释性与预测能力。
提出的方法
- 提出一种高维潜在面板分位数回归模型,其线性组合包含可观测特征与不可观测的潜在因子。
- 采用分位数损失函数,并结合ℓ₁正则化以估计稀疏系数,以及核范数正则化以估计低秩因子载荷与因子收益。
- 利用ADMM算法高效求解复合优化问题,实现可扩展估计。
- 允许存在时间依赖性与重尾误差分布,这些在金融数据中普遍存在。
- 在多个分位数(如10%、50%、90%)上联合估计模型,以捕捉异质性风险暴露。
- 引入滞后因变量,并通过潜在因子控制遗漏变量偏差。
实验结果
研究问题
- RQ1哪些公司特征在收益分布的不同分位数上具有显著的预测能力?
- RQ2资产对潜在因子的敏感性在收益分布的下、中、上分位数之间有何差异?
- RQ3在高维、依赖性与重尾数据下,联合模型能否一致估计稀疏特征与低秩潜在因子?
- RQ4潜在因子在多大程度上吸收了特征的预测能力,从而降低其边际显著性?
- RQ5与仅依赖特征或因子的模型相比,潜在因子的引入在多大程度上提升了模型对收益变异的解释能力?
主要发现
- 在控制潜在因子后,特征的预测能力显著减弱,表明特征中的大部分信息已被共同因子捕获。
- 不同分位数下的因子载荷与系数存在显著差异:上、下分位数的风险暴露与中位数水平明显不同。
- 所提出的估计量在一般条件下(包括时间依赖性与重尾误差)均能一致估计稀疏系数与低秩因子结构。
- 在大规模金融数据面板上的实证应用表明,该模型在捕捉尾部风险与分布动态方面优于传统的均值模型或单因子模型。
- 该方法成功识别出一组在极端分位数上具有显著预测能力的特征(如规模、动量、盈利能力),而其他特征则主要受潜在因子主导。
- 核范数正则化有效捕捉了潜在因子的低秩结构,即使在资产数量与时间跨度较大的情况下亦成立。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。