Skip to main content
QUICK REVIEW

[논문 리뷰] Is Medical Chest X-ray Data Anonymous?

Kai Packhäuser, Sebastian Gündel|arXiv (Cornell University)|2021. 03. 15.
COVID-19 diagnosis using AI참고 문헌 41인용 수 11
한 줄 요약

이 연구는 심층 학습 모델이 익명화된 의료 흉부 X-ray 영상에서 높은 정확도로 환자를 재식별할 수 있음을 입증한다. ChestX-ray14 데이터셋에서 AUC는 0.9940, 분류 정확도는 95.55%를 기록하였다. 연구 결과는 단순한 식별자 제거를 통한 익명화가 초기 영상 촬영 이후 수 년이 지난 후에도 환자 개인정보 보호에 부적합하다는 것을 시사한다.

ABSTRACT

With the rise and ever-increasing potential of deep learning techniques in recent years, publicly available medical datasets became a key factor to enable reproducible development of diagnostic algorithms in the medical domain. Medical data contains sensitive patient-related information and is therefore usually anonymized by removing patient identifiers, e.g., patient names before publication. To the best of our knowledge, we are the first to show that a well-trained deep learning system is able to recover the patient identity from chest X-ray data. We demonstrate this using the publicly available large-scale ChestX-ray14 dataset, a collection of 112,120 frontal-view chest X-ray images from 30,805 unique patients. Our verification system is able to identify whether two frontal chest X-ray images are from the same person with an AUC of 0.9940 and a classification accuracy of 95.55%. We further highlight that the proposed system is able to reveal the same person even ten and more years after the initial scan. When pursuing a retrieval approach, we observe an mAP@R of 0.9748 and a precision@1 of 0.9963. Furthermore, we achieve an AUC of up to 0.9870 and a precision@1 of up to 0.9444 when evaluating our trained networks on CheXpert and the COVID-19 Image Data Collection. Based on this high identification rate, a potential attacker may leak patient-related information and additionally cross-reference images to obtain more information. Thus, there is a great risk of sensitive content falling into unauthorized hands or being disseminated against the will of the concerned patients. Especially during the COVID-19 pandemic, numerous chest X-ray datasets have been published to advance research. Therefore, such data may be vulnerable to potential attacks by deep learning-based re-identification algorithms.

연구 동기 및 목표

  • 익명화된 의료 흉부 X-ray 영상이 심층 학습을 통해 개별 환자와 연결될 수 있는지 조사하기 위해.
  • 촬영 후 10년이 넘는 기간이 경과한 이후에도 재식별 성능가지의 내구성을 평가하기 위해.
  • ChestX-ray14, CheXpert, 그리고 COVID-19 Image Data Collection과 같은 공개 데이터셋이 신원 재구성에 얼마나 취약한지 평가하기 위해.
  • 특히 코로나19 패안드믹 기간 동안 의료 영상 데이터를 공개할 경우 발생할 수 있는 개인정보 유출 위험을 부각하기 위해.

제안 방법

  • 두 장의 전면 투영 흉부 X-ray 영상이 동일한 환자 소유인지 판단하기 위해 심층 학습 기반의 확인 시스템을 개발하였다.
  • 이미지 임베딩을 비교하고 신원 유사도를 예측하기 위해 시아모이 신경망 아키텍처를 사용하였다.
  • 이 시스템은 30,805명의 유니크 환자로부터 112,120장의 영상이 포함된 ChestX-ray14 데이터셋에서 평가되었다.
  • 재식별 성능를 측정하기 위해 mAP@R 및 precision@1을 사용한 검색 기반 평가를 실시하였다.
  • 외부 데이터셋인 CheXpert 및 COVID-19 Image Data Collection을 대상으로 일반화 능력을 테스트하였다.
  • 다양한 평가 프로토콜을 통해 AUC, 정확도, mAP@R, precision@1을 사용하여 성능를 정량화하였다.

실험 결과

연구 질문

  • RQ1심층 학습 모델은 익명화된 전면 흉부 X-ray 영상에서 고정확도로 환자를 재식별할 수 있는가?
  • RQ2촬영 후 10년이 넘는 시간 간격이 있는 영상 간 재식별은 얼마나 효과적인가?
  • RQ3이 모델은 다른 공개된 의료 영상 데이터셋으로 일반화될 수 있는가?
  • RQ4심층 학습의 맥락에서 익명화된 의료 영상 데이터를 공개할 경우 개인정보 보호에 어떤 위험이 존재하는가?

주요 결과

  • 확인 시스템은 ChestX-ray14 데이터셋에서 AUC 0.9940, 분류 정확도 95.55%를 기록하였다.
  • 모델은 높은 성능을 유지하며 재식별 작업을 수행하였고, mAP@R는 0.9748, precision@1은 0.9963을 달성하였다.
  • CheXpert 데이터셋에서는 AUC 최대 0.9870, precision@1 최대 0.9444를 기록하였다.
  • 촬영 후 10년이 넘는 시간 간격이 있는 영상 간에도 시스템은 환자를 성공적으로 재식별하였다.
  • 현재의 익명화 방식은 심층 학습을 통한 신원 재구성 방지를 위해 부적절하다는 것이 결과적으로 드러났다.
  • 특히 공개된 의료 영상 데이터셋을 통해 환자 개인정보 유출 위험이 상당히 존재한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.