[논문 리뷰] GasHis-Transformer: A Multi-scale Visual Transformer Approach for Gastric Histopathology Image Classification.
이 논문은 전역 및 국소 특징 추출 모듈을 통합한 다중 척도 시각 트랜스포머 모델인 GasHis-Transformer를 제안한다. 이는 공개된 H&E 염색 데이터셋에서 98.0%의 정확도, 98.0%의 정밀도, 100.0%의 재현율, 96.0%의 F1 점수를 기록하여 높은 성능과 노이즈 및 적대적 공격에 대한 강건성을 입증한다.
Existing deep learning methods for diagnosis of gastric cancer commonly use convolutional neural network. Recently, the Visual Transformer has attracted great attention because of its performance and efficiency, but its applications are mostly in the field of computer vision. In this paper, a multi-scale visual transformer model, referred to as GasHis-Transformer, is proposed for Gastric Histopathological Image Classification (GHIC), which enables the automatic classification of microscopic gastric images into abnormal and normal cases. The GasHis-Transformer model consists of two key modules: A global information module and a local information module to extract histopathological features effectively. In our experiments, a public hematoxylin and eosin (H&E) stained gastric histopathological dataset with 280 abnormal and normal images are divided into training, validation and test sets by a ratio of 1 : 1 : 2. The GasHis-Transformer model is applied to estimate precision, recall, F1-score and accuracy on the test set of gastric histopathological dataset as 98.0%, 100.0%, 96.0% and 98.0%, respectively. Furthermore, a critical study is conducted to evaluate the robustness of GasHis-Transformer, where ten different noises including four adversarial attack and six conventional image noises are added. In addition, a clinically meaningful study is executed to test the gastrointestinal cancer identification performance of GasHis-Transformer with 620 abnormal images and achieves 96.8% accuracy. Finally, a comparative study is performed to test the generalizability with both H&E and immunohistochemical stained images on a lymphoma image dataset and a breast cancer dataset, producing comparable F1-scores (85.6% and 82.8%) and accuracies (83.9% and 89.4%), respectively. In conclusion, GasHisTransformer demonstrates high classification performance and shows its significant potential in the GHIC task.
연구 동기 및 목표
- 기존의 컨volutional 신경망이 위 조직병리학 영상 분류에서 가지는 한계를 해결하기 위해 시각 트랜스포머 기반 접근 방식을 도입하기 위해.
- 다중 척도 트랜스포머 아키텍처에서 전역 및 국소 정보 모듈을 결합하여 위 조직병리학 영상의 특징 추출을 향상시키기 위해.
- 임상적 신뢰성을 확보하기 위해 다양한 영상 손상과 적대적 공격 하에서 모델의 강건성을 평가하기 위해.
- 다양한 염색 유형(H&E 및 면역형광염색)을 포함한 다양한 암 데이터셋에서 모델의 일반화 능력을 테스트하기 위해.
- 위암, 림프종, 유방암 데이터셋에서 기존 방법과 비교하거나 이를 초월하는 높은 진단 성능를 달성하기 위해.
제안 방법
- GasHis-Transformer 모델은 전역 정보 모듈과 국소 정보 모듈을 갖춘 이중 브랜치 아키텍처를 사용하여 장거리 및 세밀한 조직병리학적 특징을 모두 캡처한다.
- 입력 패치를 다양한 해상도에서 처리하여 다중 척도 특징 표현을 구현함으로써 H&E 염색 위 조직 영상에서 계층적 패턴을 학습할 수 있도록 한다.
- 시각 트랜스포머 인코더는 전통적인 컨볼루션 연산을 대체하기 위해 자기주의 메커니즘을 사용하여 이미지 패치 간의 의존성을 모델링한다.
- [CLS] 토큰 표현에 분류 헤드를 적용하여 정상 또는 비정상 위 조직 상태를 예측한다.
- 강건성 평가를 위해 10종류의 노이즈를 주입하였으며, 이는 4종류의 적대적 공격과 6종류의 전통적 영상 손상이 포함된다.
- 일반화 능력을 검증하기 위해 외부 데이터셋(림프종 및 유방암 영상 포함)을 사용하여 H&E 및 면역형광염색 프로토콜을 모두 적용한다.
실험 결과
연구 질문
- RQ1다중 척도 시각 트랜스포머 모델이 기존의 CNN보다 위 조직병리학 영상 분류에서 뛰어난 성능을 보일 수 있는가?
- RQ2전역 및 국소 특징 추출의 통합이 GHIC에서 분류 성능 향상에 어떤 영향을 미치는가?
- RQ3임상 진단 맥락에서 GasHis-Transformer는 영상 노이즈와 적대적 공격에 얼마나 강건한가?
- RQ4림프종 및 유방암과 같은 다양한 염색 기법과 암 유형 간에서 모델의 일반화 능력은 어느 정도인가?
- RQ5대규모 임상적으로 관련성이 높은 위암 데이터셋에서 GasHis-Transformer의 진단 정확도는 얼마인가?
주요 결과
- GasHis-Transformer는 H&E 염색 위 조직병리학 영상 280장으로 구성된 테스트 세트에서 98.0%의 정확도, 98.0%의 정밀도, 100.0%의 재현율, 96.0%의 F1 점수를 기록하였다.
- 모델는 10종류의 다양한 노이즈, 적대적 공격 포함하더라도 높은 성능를 유지하는 강력한 강건성을 보였다.
- 620장의 비정상 위 조직 영상으로 구성된 더 큰 임상 데이터셋에서, 모델는 위암 병변을 식별하는 데 96.8%의 정확도를 달성하였다.
- 림프종 데이터셋에서 평가한 결과, GasHis-Transformer는 H&E 및 면역형광염색 모두에서 F1 점수 85.6%와 정확도 83.9%를 기록하였다.
- 유방암 데이터셋에서 모델는 F1 점수 82.8%와 정확도 89.4%를 기록하여 다양한 암 유형과 염색 방법 간의 강력한 일반화 능력을 보였다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.