[論文レビュー] Is Medical Chest X-ray Data Anonymous?
本研究では、深層学習モデルが匿名化された胸部X線画像から高精度で患者を再識別できることを示している。ChestX-ray14データセット上でAUC 0.9940、分類精度95.55%を達成した。研究結果から、単純な識別子の削除による匿名化では、初期撮影から何年も経過した後でさえも、患者のプライバシー保護が不十分であることが明らかになった。
With the rise and ever-increasing potential of deep learning techniques in recent years, publicly available medical datasets became a key factor to enable reproducible development of diagnostic algorithms in the medical domain. Medical data contains sensitive patient-related information and is therefore usually anonymized by removing patient identifiers, e.g., patient names before publication. To the best of our knowledge, we are the first to show that a well-trained deep learning system is able to recover the patient identity from chest X-ray data. We demonstrate this using the publicly available large-scale ChestX-ray14 dataset, a collection of 112,120 frontal-view chest X-ray images from 30,805 unique patients. Our verification system is able to identify whether two frontal chest X-ray images are from the same person with an AUC of 0.9940 and a classification accuracy of 95.55%. We further highlight that the proposed system is able to reveal the same person even ten and more years after the initial scan. When pursuing a retrieval approach, we observe an mAP@R of 0.9748 and a precision@1 of 0.9963. Furthermore, we achieve an AUC of up to 0.9870 and a precision@1 of up to 0.9444 when evaluating our trained networks on CheXpert and the COVID-19 Image Data Collection. Based on this high identification rate, a potential attacker may leak patient-related information and additionally cross-reference images to obtain more information. Thus, there is a great risk of sensitive content falling into unauthorized hands or being disseminated against the will of the concerned patients. Especially during the COVID-19 pandemic, numerous chest X-ray datasets have been published to advance research. Therefore, such data may be vulnerable to potential attacks by deep learning-based re-identification algorithms.
研究の動機と目的
- 匿名化された医療用胸部X線画像が、深層学習を用いて個々の患者に再関連付け可能かどうかを調査すること。
- 撮影後10年以上経過した画像間でも再識別が可能かどうかの耐久性を評価すること。
- ChestX-ray14、CheXpert、COVID-19 Image Data Collectionといった公開データセットが、身元再構築に対してどれほど脆弱であるかを評価すること。
- 特にCOVID-19パンデミックの文脈において、医療画像データを公開することに伴うプライバシーリスクを強調すること。
提案手法
- 2枚の前向き撮影胸部X線画像が同じ患者に属するかどうかを判断するための深層学習ベースの検証システムを訓練した。
- モデルは、画像埋め込みを比較して身元類似度を予測するため、シアンプス型ニューラルネットワークアーキテクチャを用いた。
- システムは、30,805人の異なる患者からなる112,120枚の画像を含むChestX-ray14データセットで評価された。
- 再識別性能を測定するために、mAP@Rおよびprecision@1を用いたリtrievalベースの評価が実施された。
- 外部データセット(CheXpertおよびCOVID-19 Image Data Collection)を用いた汎化性のテストが行われた。
- 複数の評価プロトコルにわたり、AUC、精度、mAP@R、precision@1を用いて性能を定量化した。
実験結果
リサーチクエスチョン
- RQ1深層学習モデルは、匿名化された前向き撮影胸部X線画像から高精度で患者を再識別可能か?
- RQ2撮影後10年以上離れた画像間でも再識別はどの程度効果的か?
- RQ3このモデルは、他の公開利用可能な医療画像データセットへどの程度一般化可能か?
- RQ4深層学習の文脈において、匿名化された医療画像データを公開することのプライバシー的影響は何か?
主な発見
- 検証システムはChestX-ray14データセットでAUC 0.9940、分類精度95.55%を達成した。
- モデルは再識別タスクにおいて高い性能を維持し、mAP@R 0.9748、precision@1 0.9963を達成した。
- CheXpertデータセットでは、AUCが最大0.9870、precision@1が最大0.9444に達した。
- 画像が10年以上離れて撮影された場合でも、システムは患者を正常に再識別できた。
- 現在の匿名化手法では、深層学習を用いた身元再構築を防止できないことが示唆された。
- 特に公開共有された医療画像データセットにおいて、患者のプライバシー侵害のリスクが顕著に存在する。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。