Skip to main content
QUICK REVIEW

[논문 리뷰] Data aggregation can lead to biased inferences in Bayesian linear mixed models

Daniel J. Schad, Bruno Nicenboim|Data Archiving and Networked Services (DANS)|2022. 01. 01.
Psychometric Methodologies and Testing인용 수 5
한 줄 요약

이 연구는 구형성 가정이 위반되거나 항목 수준의 분산이 忽시될 경우 베이지안 선형 혼합 모델에서 데이터 집계가 편향된 베이즈 요인을 초래함을 보여준다. 시뮬레이션 기반 校정을 통해 저자들은 집계된 분석이 과도하게 유리하거나 보수적인 증거를 초래하는 반면, 전체 무작위 효과 구조를 가진 비집약 모델은 편향 없는 추론을 제공함을 밝혀냈다.

ABSTRACT

Bayesian linear mixed-effects models are increasingly being used in the cognitive sciences to perform null hypothesis tests, where a null hypothesis that an effect is zero is compared with an alternative hypothesis that the effect exists and is different from zero. While software tools for Bayes factor null hypothesis tests are easily accessible, how to specify the data and the model correctly is often not clear. In Bayesian approaches, many authors recommend data aggregation at the by-subject level and running Bayes factors on aggregated data. Here, we use simulation-based calibration for model inference to demonstrate that null hypothesis tests can yield biased Bayes factors, when computed from aggregated data. Specifically, when random slope variances differ (i.e., violated sphericity assumption), Bayes factors are too conservative for contrasts where the variance is small and they are too liberal for contrasts where the variance is large. Moreover, Bayes factors for by-subject aggregated data are biased (too liberal) when random item variance is present but ignored in the analysis. We also perform corresponding frequentist analyses (type I and II error probabilities) to illustrate that the same problems exist and are well known from frequentist tools. These problems can be circumvented by running Bayesian linear mixed-effects models on non-aggregated data such as on individual trials and by explicitly modeling the full random effects structure. Reproducible code is available from https://osf.io/mjf47/.

연구 동기 및 목표

  • 베이지안 선형 혼합 모델에서 데이터 집계가 편향된 베이즈 요인 추론을 유도하는지 조사하기.
  • 집계된 자료에서 베이즈 요인 신뢰성에 영향을 미치는 구형성 가정 위반의 영향을 검토하기.
  • 참고자 수준의 분산을 모형화하지 않을 경우, 집계된 분석에서 베이즈 요인 정확도에 미치는 영향 평가하기.
  • 진짜 효과를 탐지하는 데 있어 집계된 모델과 비집약된 모델의 성능 비교하기.
  • 베이지안 반복 측정 설계에서 무작위 효과 구조를 지정하는 데 있어 최선의 실천 방법 제공하기.

제안 방법

  • 후행 예측 p-값과 베이즈 요인의 신뢰성을 평가하기 위해 시뮬레이션 기반 校정(SBC)을 시행함.
  • 구형성 위반을 시험하기 위해 다양한 무작위 기울기 분산과 항목 수준의 분산을 가진 다수의 실험 설계를 시뮬레이션함.
  • 집계된 자료와 비집약된 자료 모두에 대해 brms와 BayesFactor 패키지를 사용해 베이지안 선형 혼합 모델(LMM)과 ANOVA를 적용함.
  • 계층 모델에서의 모델 비교를 위해 bridgesampling을 사용해 베이즈 요인을 계산함.
  • 조건별로 기대된 후행 예측 p-값과 관측된 후행 예측 p-값을 비교함으로써 베이즈 요인의 편향 평가함.
  • 전체 모델 대비 감소된 모델을 포함한 다양한 무작위 효과 구조에서의 모델 성능 평가함.

실험 결과

연구 질문

  • RQ1데이터를 참조자 수준으로 집계하는 것이 베이지안 선형 혼합 모델에서 베이즈 요인 추정치에 편향을 유도하는가?
  • RQ2구형성 가정 위반이 집계된 자료 분석에서 베이즈 요인 정확도에 어떤 영향을 미치는가?
  • RQ3데이터를 집계할 경우 항목 수준의 분산을 모형화하지 않을 경우 베이즈 요인 추론에 어떤 영향을 미치는가?
  • RQ4비집약된 베이지안 모델에서 전체 무작위 효과 구조를 사용하면 베이즈 요인 추정의 편향을 제거할 수 있는가?
  • RQ5어떤 조건에서 조차도 집계된 자료를 사용할 경우 베이즈 요인이 편향되지 않는가?

주요 결과

  • 구형성이 위반될 경우, 큰 무작위 기울기 분산을 가진 대비에서는 베이즈 요인이 과도하게 유리해지고, 작은 분산을 가진 대비에서는 과도하게 보수적으로 나타남.
  • 항목 수준의 분산을 忽시한 채 데이터를 집계하면 효과에 대한 증거를 과대평가하는 편향된 베이즈 요인이 초래됨.
  • 참조자와 항목의 양방향 무작위 기울기 포함한 비집약 모델은 베이즈 요인 추정의 편향을 상당히 감소시키거나 제거함.
  • 조금의 편향이 여전히 존재함. 특히 구형성이 위반되었을 경우, BayesFactor 패키지 분석에서 비집약된 자료를 사용하더라도 편향이 발생함으로써 기본 모델 사양의 한계를 시사함.
  • 베이지안 ANOVA에서 집계된 자료를 사용할 경우, 구형성 가정이 충족되지 않으면 체계적으로 왜곡된 증거가 초래됨.
  • 베이지안 반복 측정 설계에서 편향 없는 추론을 위해서는 비집약된 자료에서 참조자와 항목의 양방향 무작위 효과를 명시적으로 모형화하는 것이 필수적임.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.