[论文解读] Inferred vs traditional personality assessment: are we predicting the same thing?
本研究评估了从数字足迹预测人格的机器学习模型是否能产生与传统自评评估相当的结果。分析220篇论文后发现,预测的人格特质与自评特质之间的相关性仅中等程度(r ≈ 0.42–0.48),未达到标准心理测量信度水平,且缺乏稳定性和维度结构,表明推断出的人格特质并不等同于既定的人格构念。
Machine learning methods are widely used by researchers to predict psychological characteristics from digital records. To find out whether automatic personality estimates retain the properties of the original traits, we reviewed 220 recent articles. First, we put together the predictive quality estimates from a subset of the studies which declare separation of training, validation, and testing phases, which is critical for ensuring the correctness of quality estimates in machine learning. Only 20% of the reviewed papers met this criterion. To compare the reported quality estimates, we converted them to approximate Pearson correlations. The credible upper limits for correlations between predicted and self-reported personality traits vary in a range between 0.42 and 0.48, depending on the specific trait. The achieved values are substantially below the correlations between traits measured with distinct self-report questionnaires. This suggests that we cannot readily interpret personality predictions as estimates of the original traits or expect predicted personality traits to reproduce known relationships with life outcomes regularly. Next, we complement quality estimates evaluation with evidence on psychometric properties of predicted traits. The few existing results suggest that predicted traits are less stable with time and have lower effective dimensionality than self-reported personality. The predictive text-based models perform substantially worse outside their training domains but stay above a random baseline. The evidence on the relationships between predicted traits and external variables is mixed. Predictive features are difficult to use for validation, due to the lack of prior hypotheses. Thus, predicted personality traits fail to retain important properties of the original characteristics. This calls for the cautious use and targeted validation of the predictive models.
研究动机与目标
- 确定从数字足迹推断人格的机器学习模型是否能产生与传统自评人格评估相当的结果。
- 评估推断人格特质的心理测量质量,包括信度、稳定性及有效维度数。
- 评估预测的人格特质是否保持与生活结果及外部变量之间的已知关系。
- 识别当前研究中的方法论缺陷,特别是缺乏适当的训练-验证-测试数据分离。
- 倡导在心理研究和应用中谨慎使用人格预测模型,并进行针对性验证。
提出的方法
- 系统性回顾了220篇关于从数字记录自动预测人格的近期研究。
- 仅选取报告了独立训练、验证和测试阶段的研究,以确保性能估计的有效性。
- 将报告的性能指标(如AUC、F1、R²)转换为近似皮尔逊相关系数,以实现跨研究比较。
- 利用现有证据评估预测特质的心理测量属性,包括重测信度、时间稳定性及有效维度数。
- 通过分析预测特质与生活结果或行为变量之间报告的关系,评估外部效度。
- 采用心理测量工具验证框架,将推断特质的质量与既定标准进行对比评估。
实验结果
研究问题
- RQ1从数字足迹预测人格的机器学习模型与自评人格特质之间的相关性在多大程度上成立?
- RQ2与传统自评测量相比,预测的人格特质在时间上是否表现出足够的信度和稳定性?
- RQ3推断的人格特质是否保持与大五人格模型相同的有效维度数和结构完整性?
- RQ4预测的人格特质是否与外部变量(如生活结果或行为模式)存在有意义的关系?
- RQ5在模型评估方面存在哪些方法论缺陷,特别是如何削弱了当前人格预测研究的信心?
主要发现
- 预测特质与自评特质之间相关性的可信上限范围为0.42至0.48,远低于不同自评量表之间通常观察到的0.6–0.9相关性。
- 仅20%的被审查研究正确分离了训练、验证和测试数据,引发了对其报告性能估计有效性的担忧。
- 与自评特质相比,预测的人格特质表现出较低的重测信度和较弱的时间稳定性。
- 推断特质的有效维度数低于传统人格构念,表明其结构完整性受损。
- 基于文本的模型在训练领域之外性能显著下降,尽管仍高于随机基线水平。
- 关于预测特质与外部变量之间关系的证据不一致,未观察到强烈或可重复的模式。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。