Skip to main content
QUICK REVIEW

[논문 리뷰] Model-assisted cohort selection with bias analysis for generating large-scale cohorts from the EHR for oncology research

Benjamin Birnbaum, Nathan C. Nussbaum|arXiv (Cornell University)|2020. 01. 13.
Machine Learning in Healthcare참고 문헌 11인용 수 200
한 줄 요약

요약: 이 논문은 Bias Analysis를 포함한 Model-Assisted Cohort Selection (MACS)를 도입하여 EHR 기반 종양학 코호트를 대규모로 효율적으로 생성하고, 예측 성능이 높으며 이후 분석에서 감지 가능한 편향이 없음을 보여준다.

ABSTRACT

Objective Electronic health records (EHRs) are a promising source of data for health outcomes research in oncology. A challenge in using EHR data is that selecting cohorts of patients often requires information in unstructured parts of the record. Machine learning has been used to address this, but even high-performing algorithms may select patients in a non-random manner and bias the resulting cohort. To improve the efficiency of cohort selection while measuring potential bias, we introduce a technique called Model-Assisted Cohort Selection (MACS) with Bias Analysis and apply it to the selection of metastatic breast cancer (mBC) patients. Materials and Methods We trained a model on 17,263 patients using term-frequency inverse-document-frequency (TF-IDF) and logistic regression. We used a test set of 17,292 patients to measure algorithm performance and perform Bias Analysis. We compared the cohort generated by MACS to the cohort that would have been generated without MACS as reference standard, first by comparing distributions of an extensive set of clinical and demographic variables and then by comparing the results of two analyses addressing existing example research questions. Results Our algorithm had an area under the curve (AUC) of 0.976, a sensitivity of 96.0%, and an abstraction efficiency gain of 77.9%. During Bias Analysis, we found no large differences in baseline characteristics and no differences in the example analyses. Conclusion MACS with bias analysis can significantly improve the efficiency of cohort selection on EHR data while instilling confidence that outcomes research performed on the resulting cohort will not be biased.

연구 동기 및 목표

  • 비정형 데이터로 인한 비무작위 코호트 선택 문제를 해결하고 종양학 결과 연구를 위해 EHR 데이터를 활용하도록 동기를 부여한다.
  • 다운스트림 분석을 신뢰할 수 있도록 편향 평가를 포함하는 확장 가능한 코호트 선택 방법을 개발한다.
  • 전이성 유방암(mBC)에 이 방법을 적용하여 효율성과 편향 억제를 입증한다.

제안 방법

  • 대상 코호트를 식별하기 위해 17,263명의 환자에서 TF-IDF + 로지스틱 회귀 모델을 훈련한다.
  • 17,292명의 분리된 테스트 세트에서 성능을 평가한다.
  • 다수의 임상 및 인구통계 변수에 걸쳐 MACS로 생성된 코호트를 기준 표준과 비교하는 편향 분석을 수행한다.
  • MACS 코호트와 비-MACS 코호트 간 변수 분포를 비교한다.
  • (i) MACS가 높은 판별력을 달성하고 (ii) 편향이 예시 분석에 실질적으로 영향을 미치지 않음을 보여준다.

실험 결과

연구 질문

  • RQ1MACS가 EHR 데이터에서 종양학 연구를 위한 코호트 선별의 효율성을 향상시킬 수 있는가?
  • RQ2참고 표준과 비교했을 때 MACS가 기저 특성에 대해 검출 가능한 편향을 도입하는가?
  • RQ3MACS 생성 코호트에서 수행된 분석이 편향이 없는 기준 분석과 일치하는 결과를 내는가?

주요 결과

  • MACS 선택자의 AUC 0.976은 강한 판별력을 나타낸다.
  • 민감도 96.0%는 대상 코호트의 높은 진양성 포착을 보여준다.
  • 추상화 효율성 증가 77.9%는 상당한 워크플로우 개선을 보여준다.
  • 편향 분석에서 MACS와 기준 코호트 간 기저 특성에 큰 차이가 없음이 나타났다.
  • MACS 유래 코호트와 기준 분석 간의 예시 분석에서 차이가 발견되지 않았다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.