[논문 리뷰] Convolution Neural Networks for diagnosing colon and lung cancer histopathological images
A shallow CNN is trained on the LC25000 histopathology dataset to classify lung and colon cancer images, achieving around 98% accuracy on both tasks without heavy preprocessing. The study analyzes architecture, training strategy, and results.
Lung and Colon cancer are one of the leading causes of mortality and morbidity in adults. Histopathological diagnosis is one of the key components to discern cancer type. The aim of the present research is to propose a computer aided diagnosis system for diagnosing squamous cell carcinomas and adenocarcinomas of lung as well as adenocarcinomas of colon using convolutional neural networks by evaluating the digital pathology images for these cancers. Hereby, rendering artificial intelligence as useful technology in the near future. A total of 2500 digital images were acquired from LC25000 dataset containing 5000 images for each class. A shallow neural network architecture was used classify the histopathological slides into squamous cell carcinomas, adenocarcinomas and benign for the lung. Similar model was used to classify adenocarcinomas and benign for colon. The diagnostic accuracy of more than 97% and 96% was recorded for lung and colon respectively.
연구 동기 및 목표
- 폐 및 대장 암의 조직병리학 이미지에서 컴퓨터 보조 진단을 가능하게 하는 동기를 부여한다.
- 얕은 CNN이 폐의 편평상피세포암과 선종, 대장의 선종을 효과적으로 분류할 수 있는지 평가한다.
- 고해상도 조직병리학 이미지를 과도한 다운샘플링 없이 학습할 수 있는지 보여준다.
- 재현을 위한 학습/평가 프로토콜 및 모델 아티팩트를 제공한다.
제안 방법
- 클래스당 5,000장의 이미지로 구성된 5개 클래스의 LC25000 데이터셋을 augmentation으로 확장하여 25,000장으로 사용한다.
- 3개의 합성곱 층 CNN(필터 수 32, 64, 64, 커널 3x3, 스트라이드 2) 뒤에 두 개의 완전 연결 층(512 및 3 또는 2 출력)을 구성한다.
- 각 합성곱 층 뒤에 ReLU 활성화와 최대 풀링, Dense 층 사이에 드롭아웃을 적용한다.
- RMSprop으로 학습하고, 미니 배치 크기 32, 학습률 1e-4, rho 0.9, epsilon 1e-7, 손실 함수는 범주형 교차 엔트로피를 사용하여 100 반복 학습한다.
- 80-10-10의 학습/검증/테스트 분할에서 이미지 수준 정확도와 손실을 보고하고, TensorFlow로 모델을 구현하며 h5 아티팩트를 공유한다.
실험 결과
연구 질문
- RQ1얕은 CNN이 폐 조직병리학 이미지를 선종(adeno), 편평세포암(squamous cell carcinoma) 또는 양성(benign)으로 정확하게 분류할 수 있는가?
- RQ2동일한 아키텍처가 대장 조직병리학 이미지를 선종(adeno) 또는 양성으로 높은 정확도로 분류할 수 있는가?
- RQ3LC25000에서 학습된 얕은 CNN이 동일한 데이터에 대해 전통적인 ML 방법 및 전이 학습 접근법을 능가하는가?
주요 결과
- 폐 모델은 학습 정확도 97.9216% 및 손실 5.83230을 달성했고, 검증 정확도는 97.8987% 및 손실 6.11450%를 달성했다.
- 대장 모델은 학습 정확도 96.9503% 및 손실 7.9340%를 달성했고, 검증 정확도는 96.6110% 및 손실 9.7141%를 달성했다.
- 두 모델은 대략 20 에폭에서 수렴했고, 드롭아웃은 정확도/손실 그래프에 변동성을 도입했지만 일반화에 도움을 주었다.
- 이 접근법은 얕은 CNN이 고해상도 조직병리학 이미지를 과도한 다운샘플링 없이 효과적으로 분류할 수 있음을 시연한다.
- 특징 맵은 학습된 필터가 가브르(Gabor) 필터와 유사한 질감-like 패턴을 닮고 있어 학습된 표현을 뒷받침한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.