Skip to main content
QUICK REVIEW

[논문 리뷰] Workflow Techniques for the Robust Use of Bayes Factors

Daniel J. Schad, Bruno Nicenboim|arXiv (Cornell University)|2021. 03. 15.
Explainable Artificial Intelligence (XAI)인용 수 6
한 줄 요약

이 논문은 인지과학 연구에서 베이즈 요인의 탄력성을 평가하기 위한 체계적인 워크플로우를 제안한다. 이는 사전 분포 민감도, 추정 불안정성, 데이터 변동성, 그리고 불확실성 하에서의 의사결정 문제를 다룬다. 시뮬레이션 기반 校정과 민감도 분석을 통해 저자들은 베이즈 요인이 나쁜 사전 분포 설정, 부족한 유효 표본 크기, 모형 오설정으로 인해 신뢰할 수 없을 수 있음을 입증하며, 유틸리티 기반 의사결정과 재현 가능한 校정을 통해 실무에서 타당한 추론을 확보할 것을 주장한다.

ABSTRACT

Inferences about hypotheses are ubiquitous in the cognitive sciences. Bayes factors provide one general way to compare different hypotheses by their compatibility with the observed data. Those quantifications can then also be used to choose between hypotheses. While Bayes factors provide an immediate approach to hypothesis testing, they are highly sensitive to details of the data/model assumptions. Moreover it's not clear how straightforwardly this approach can be implemented in practice, and in particular how sensitive it is to the details of the computational implementation. Here, we investigate these questions for Bayes factor analyses in the cognitive sciences. We explain the statistics underlying Bayes factors as a tool for Bayesian inferences and discuss that utility functions are needed for principled decisions on hypotheses. Next, we study how Bayes factors misbehave under different conditions. This includes a study of errors in the estimation of Bayes factors. Importantly, it is unknown whether Bayes factor estimates based on bridge sampling are unbiased for complex analyses. We are the first to use simulation-based calibration as a tool to test the accuracy of Bayes factor estimates. Moreover, we study how stable Bayes factors are against different MCMC draws. We moreover study how Bayes factors depend on variation in the data. We also look at variability of decisions based on Bayes factors and how to optimize decisions using a utility function. We outline a Bayes factor workflow that researchers can use to study whether Bayes factors are robust for their individual analysis, and we illustrate this workflow using an example from the cognitive sciences. We hope that this study will provide a workflow to test the strengths and limitations of Bayes factors as a way to quantify evidence in support of scientific hypotheses. Reproducible code is available from https://osf.io/y354c/.

연구 동기 및 목표

  • 다양한 데이터, 모형, 사전 분포 가정 하에서 인지과학 적용에서 베이즈 요인의 탄력성을 조사하기 위해.
  • 베이즈 요인 추정의 불안정성 원인을 특정화하기 위해, 특히 부적절한 사전 분포 설정과 부족한 MCMC 유효 표본 크기의 영향을 분석하기 위해.
  • 시뮬레이션 기반 校정(SBC)을 사용하여 다리 샘플링과 세이버지-딕키 방법 추정치의 정확도를 평가하기 위해.
  • 반복적인 데이터 추출과 복제 시도에서 베이즈 요인 결과의 변동성을 분석하기 위해.
  • 연구자들이 유틸리티 함수와 시뮬레이션 기반 校정을 사용하여 베이즈 요인 추론과 의사결정의 신뢰성을 평가할 수 있는 실용적 워크플로우를 개발하기 위해.

제안 방법

  • 정확한 참 모형이 알려진 조건 하에서 시뮬레이션된 데이터로부터 사전 분포를 복원함으로써, 베이즈 요인 추정치의 정확도를 시험하기 위해 시뮬레이션 기반 校정(SBC)을 활용한다.
  • 복잡한 인지과학 데이터에 대해 계층 베이지안 모형을 피팅하고, 브릿지 샘플링을 통해 베이즈 요인을 추정하기 위해 R 패키지 brms를 사용한다.
  • 사전 분포를 다양하게 변화시켜 베이즈 요인의 안정성과 결론에 미치는 영향을 평가하기 위해 사전 민감도 분 析를 수행한다.
  • 모형 가정과 데이터의 호환성을 평가하기 위해 사전 예측 및 사후 예측 점검을 실시한다.
  • 의사결정을 체계화하고 불확실성 하에서 결론의 탄력성을 평가하기 위해 유틸리티 함수를 적용한다.
  • SBC 시뮬레이션을 통해 의사결정을 校정하고, 사후 모형 확률이 참 모형 빈도와 일치하는지 평가한다.
Figure 1: Shown are the schematic relations between the data and the model, Bayes factors, and resulting inferences and decisions. The data and the model constitute a true Bayes factor, that can be used for data informed inferences and decisions (dark red arrows). However, the true Bayes factor is u
Figure 1: Shown are the schematic relations between the data and the model, Bayes factors, and resulting inferences and decisions. The data and the model constitute a true Bayes factor, that can be used for data informed inferences and decisions (dark red arrows). However, the true Bayes factor is u

실험 결과

연구 질문

  • RQ1브릿지 샘플링을 통해 확보된 베이즈 요인 추정치는 얼마나 정확한가? 어떤 조건에서 참 베이즈 요인을 회복하지 못하는가?
  • RQ2특히 소형에서 중형 표본 설정에서, 베이즈 요인은 사전 분포의 변화에 얼마나 민감한가?
  • RQ3데이터 변동성(예: 피험자 및 자료 효과)으로 인해 반복적 표본 추출에서 베이즈 요인 결과가 얼마나 변동하는가?
  • RQ4다양한 MCMC 추출에서 베이즈 요인 추정치는 얼마나 안정적인가? 신뢰할 수 있는 추정을 위해 필요한 유효 표본 크기는 얼마인가?
  • RQ5시뮬레이션 기반 校정은 모형 오설정이나 부적절한 사전 선택으로 인한 베이즈 요인 추정치의 편향을 감지할 수 있는가?

주요 결과

  • 모형 오설정이 있을 경우, 심지어 큰 MCMC 표본 크기를 가질지라도, 브릿지 샘플링을 통한 베이즈 요인 추정치는 정확도와 안정성에 결함이 있을 수 있다.
  • SBC 결과에 따르면, 일부 모형 구성에서는 세이버지-딕키 방법을 사용할 경우 평균 사후 모형 확률이 잘못되었음을 드러내어 추론의 신뢰성이 떨어짐을 시사한다.
  • 베이즈 요인은 사전 가정에 매우 민감하며, 효과 크기 사전 분포를 변경하더라도 결과가 크게 달라지며, 이는 약한 정보 사전 범위 내에서도 마찬가지다.
  • 복제 연구 결과에 따르면, 반복적인 표본 추출에서 베이즈 요인은 넓은 범위로 변동함을 보여, 소형에서 중형 효과 크기에서는 재현성이 낮음을 시사한다.
  • 인지과학에서 흔한 소형 효과 크기나 낮은 검정력 조건에서는, 표본 크기가 크더라도 효과 크기가 상당히 클 경우를 제외하고는 베이즈 요인의 해석이 모호해진다.
  • 유틸리티 기반 의사결정은 탄력성을 크게 향상시키며, 시뮬레이션 기반 방법을 통한 校정이 이루어져야만 의사결정이 최적화됨을 보여준다.
Figure 2: Illustration of different types of parameters for two parameters Theta 1 and Theta 2. (a) Point hypothesis in all parameters. (b) Point hypothesis in some parameters. (c) Interval hypothesis in all parameters. (d) Interval hypothesis in some parameters. (e) Full hypothesis.
Figure 2: Illustration of different types of parameters for two parameters Theta 1 and Theta 2. (a) Point hypothesis in all parameters. (b) Point hypothesis in some parameters. (c) Interval hypothesis in all parameters. (d) Interval hypothesis in some parameters. (e) Full hypothesis.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.