Skip to main content
QUICK REVIEW

[논문 리뷰] Inferred vs traditional personality assessment: are we predicting the same thing?

Pavel Novikov, Лариса Марарица|arXiv (Cornell University)|2021. 03. 17.
Personality Traits and Psychology참고 문헌 162인용 수 10
한 줄 요약

이 연구는 디지털 풋프린트에서 성격을 예측하는 기계학습 모델이 전통적인 자기 보고 평가와 유사한 결과를 내는지 평가한다. 220편의 논문을 분석한 결과, 예측된 성격 특성과 자기 보고 특성 간 상관관계는 중간 수준에 머무르며(r ≈ 0.42–0.48), 표준 심리측정 신뢰도에 못 미치며, 안정성과 차원성도 부족하여 유추된 특성이 확립된 성격 구조와 동일시할 수 없다는 것을 시사한다.

ABSTRACT

Machine learning methods are widely used by researchers to predict psychological characteristics from digital records. To find out whether automatic personality estimates retain the properties of the original traits, we reviewed 220 recent articles. First, we put together the predictive quality estimates from a subset of the studies which declare separation of training, validation, and testing phases, which is critical for ensuring the correctness of quality estimates in machine learning. Only 20% of the reviewed papers met this criterion. To compare the reported quality estimates, we converted them to approximate Pearson correlations. The credible upper limits for correlations between predicted and self-reported personality traits vary in a range between 0.42 and 0.48, depending on the specific trait. The achieved values are substantially below the correlations between traits measured with distinct self-report questionnaires. This suggests that we cannot readily interpret personality predictions as estimates of the original traits or expect predicted personality traits to reproduce known relationships with life outcomes regularly. Next, we complement quality estimates evaluation with evidence on psychometric properties of predicted traits. The few existing results suggest that predicted traits are less stable with time and have lower effective dimensionality than self-reported personality. The predictive text-based models perform substantially worse outside their training domains but stay above a random baseline. The evidence on the relationships between predicted traits and external variables is mixed. Predictive features are difficult to use for validation, due to the lack of prior hypotheses. Thus, predicted personality traits fail to retain important properties of the original characteristics. This calls for the cautious use and targeted validation of the predictive models.

연구 동기 및 목표

  • 디지털 풋프린트에서 성격을 유추하는 기계학습 모델이 전통적인 자기 보고 성격 평가와 동일한 결과를 내는지 확인하기.
  • 유추된 성격 특성의 심리측정 품질, 즉 신뢰도, 안정성, 효과적 차원성 등을 평가하기.
  • 예측된 특성이 삶의 결과나 외부 변수와 기존에 알려진 관계를 유지하는지 평가하기.
  • 특히 학습-검증-테스트 분리가 부적절한 점을 포함한 현재 연구의 방법론적 한계를 규명하기.
  • 심리학 연구 및 적용 분야에서 성격 예측 모델의 신중한 사용과 목표 지향적 검증을 촉구하기.

제안 방법

  • 디지털 기록에서 자동 성격 예측에 관한 최근 220편의 연구를 체계적으로 검토.
  • 유효한 성능 추정을 보장하기 위해 학습, 검증, 테스트 단계를 별도로 보고한 연구만 선별.
  • 다양한 연구 간 비교를 위해 보고된 성능 지표(AUC, F1, R² 등)를 약간의 근사치로 피어슨 상관관계로 변환.
  • 사용 가능한 증거를 바탕으로 유추된 성격 특성의 심리측정적 성질을 평가하기 위해 재측정 신뢰도, 시간적 안정성, 효과적 차원성 등을 분석.
  • 외부 타당성을 분석하기 위해 예측된 특성과 삶의 결과 또는 행동 변수 간의 관계를 검토.
  • 기존 기준에 기반한 심리측정 도구 검증 프레임워크를 사용하여 유추된 특성의 품질을 평가.

실험 결과

연구 질문

  • RQ1디지털 풋프린트에서 성격을 예측하는 기계학습 모델이 자기 보고 성격 특성과 얼마나 상관관계가 있는가?
  • RQ2예측된 성격 특성이 전통적인 자기 보고 측정과 비교해 충분한 신뢰도와 시간적 안정성을 보이는가?
  • RQ3유추된 성격 특성이 빅파이브 모델과 동일한 효과적 차원성과 구조적 통합성을 유지하는가?
  • RQ4예측된 특성이 삶의 결과나 행동 패tern과 같은 외부 변수와 의미 있는 관련성을 가지는가?
  • RQ5특히 모델 평가 측면에서 현재 성격 예측 연구의 신뢰도를 떨어뜨리는 방법론적 결함는 무엇인가?

주요 결과

  • 유추된 성격 특성과 자기 보고 성격 특성 간 상관관계의 신뢰할 수 있는 상한선은 0.42에서 0.48 사이이며, 이는 서로 다른 자기 보고 도구 간에 일반적으로 관찰되는 0.6–0.9의 상관관계에 비해 상당히 낮다.
  • 검토된 연구 중 20%만이 학습, 검증, 테스트 데이터를 올바르게 분리하여, 보고된 성능 추정의 타당성에 우려를 제기한다.
  • 유추된 성격 특성은 자기 보고 특성에 비해 낮은 재측정 신뢰도와 감소한 시간적 안정성을 보인다.
  • 유추된 특성은 전통적인 성격 구조보다 낮은 효과적 차원성을 보이며, 이는 구조적 통합성의 손실을 시사한다.
  • 텍스트 기반 모델의 성능은 훈련 도메인 외부에서는 크게 떨어지지만, 여전히 무작위 기준 이상의 성능을 유지한다.
  • 예측된 특성과 외부 변수 간의 관계에 대한 증거는 일관성이 없으며, 강하거나 재현 가능한 패tern이 관찰되지 않는다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.