[논문 리뷰] On the uses and abuses of regression models: a call for reform of statistical practice and teaching
이 논문은 회귀 모형을 '진짜 모형'을 추구하는 것에서 연구 목적에 맞는 프레임워크로 전환함으로써 통계적 실무와 교육의 근본적인 개혁을 주장한다. 이는 서술적, 예측적, 인과적 세 가지 연구 질문 중심으로 구성된다. 논문은 목적에 따라 회귀 분석을 다르게 적용해야 하며, 공변량을 잘못 조정하거나 무작위로 모형을 피팅하는 등의 일반적인 오용을 줄일 수 있다고 주장한다.
Regression methods dominate the practice of biostatistical analysis, but biostatistical training emphasises the details of regression models and methods ahead of the purposes for which such modelling might be useful. More broadly, statistics is widely understood to provide a body of techniques for "modelling data", underpinned by what we describe as the "true model myth": that the task of the statistician/data analyst is to build a model that closely approximates the true data generating process. By way of our own historical examples and a brief review of mainstream clinical research journals, we describe how this perspective has led to a range of problems in the application of regression methods, including misguided "adjustment" for covariates, misinterpretation of regression coefficients and the widespread fitting of regression models without a clear purpose. We then outline a new approach to the teaching and application of biostatistical methods, which situates them within a framework that first requires clear definition of the substantive research question at hand within one of three categories: descriptive, predictive, or causal. Within this approach, the development and application of (multivariable) regression models, as well as other advanced biostatistical methods, should proceed differently according to the type of question. Regression methods will no doubt remain central to statistical practice as they provide a powerful tool for representing variation in a response or outcome variable as a function of "input" variables, but their conceptualisation and usage should follow from the purpose at hand.
연구 동기 및 목표
- 모형 피팅에 대한 과도한 강조로 인해 생물통계학 연구에서 회귀 모형이 널리 오용되고 있는 문제를 다루기.
- '진짜 모형의 신화'—통계학자들이 실제 데이터 생성 과정을 근사해야 한다는 믿음—를 도전하기.
- 부적절한 공변량 조정과 회귀 계수의 오해적 해석 등 임상 연구에서의 체계적 문제를 부각하기.
- 연구 질문의 유형에 맞게 회귀 분석을 적용하고 가르치는 새로운 프레임워크를 제안하기.
- 연구자가 먼저 분석 목적이 서술적, 예측적, 인과적 중 어느 것인지 명확히 하도록 요구함으로써 무작위 모형 피팅을 줄이기.
제안 방법
- 모든 통계적 연구 질문을 서술적, 예측적, 인과적 세 가지 유형으로 분류하기.
- 회귀 모형 분석을 질문 유형에 따라 설계하고 해석하는 도구로 재정의하기.
- 교육과 실무에서 '모형 중심'에서 '목적 중심'으로의 전환을 주장하기.
- 역사적 사례 분석과 주요 임상 저널의 분석을 통해 회귀 분석의 일반적인 오용 사례를 제시하기.
- 다변량 회귀 분석이 연구 목표에 따라 동일하게 적용되어서는 안 되며, 목적에 맞게 적응시켜야 한다고 강조하기.
- 연구 질문이 명확해진 후에야 통계적 명시가 이루어지는 개념적 프레임워크를 도입하기.
실험 결과
연구 질문
- RQ1왜 회귀 모형은 생물통계학 연구에서 자주 오해와 오용을 낳는가?
- RQ2'진짜 모형의 신화'는 임상 및 역학 연구에서 잘못된 통계적 실무를 어떻게 초래하는가?
- RQ3명확한 연구 목적 없이 회귀 모형을 피팅할 경우 어떤 결과가 초래되는가?
- RQ4어떻게 하면 서술적, 예측적, 인과적 연구 질문에 더 효과적으로 대응할 수 있도록 회귀 방법을 재구성할 수 있는가?
- RQ5회귀 모형 분석의 일반적인 오용을 줄이기 위해 통계 교육과 실무에 어떤 변화가 필요한가?
주요 결과
- '진짜 모형의 신화'는 공변량을 잘못 조정하거나 계수를 오해하는 등 회귀 분석의 광범위한 오용을 초래한다.
- 많은 회귀 모형이 명확한 연구 목적 없이 피팅되어 해석 가능한 가치가 없는 분석을 낳는다.
- 관찰 연구에서 잘못된 공변량 조정은 인과적 추론을 왜곡하고 타당성을 떨어뜨리는 경우가 흔하다.
- 회귀 모형이 종종 목적을 가진 도구가 아니라 목적 자체로 여겨지는 경향이 있다.
- 목적 중심의 프레임워크는 통계적 오용의 위험을 크게 줄이며 결과의 관련성과 해석 가능성까지 향상시킨다.
- 서술적, 예측적, 인과적 질문 중심으로 교육과 실무를 재구성하면 더 투명하고 타당하며 과학적으로 의미 있는 분석이 가능해진다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.