Skip to main content
QUICK REVIEW

[논문 리뷰] Benchmarking features from different radiomics toolkits / toolboxes using Image Biomarkers Standardization Initiative

Mingxi Lei, Bino Varghese|arXiv (Cornell University)|2020. 06. 23.
Radiomics and Machine Learning in Medical Imaging참고 문헌 59인용 수 5
한 줄 요약

이 연구는 알려진 기준값을 가진 디지털 페phantom을 사용하여 173개의 IBSI 표준화된 방사학적 특징을 여섯 개의 공공 툴킷과 하나의 내부 파이프라인에서 평가하였다. 이는 소프트웨어 간 상당한 차이를 드러내었으며, 특히 형태학적 특징과 회색 수준 이산화 방법에서 두드러졌고, 재현성과 일반화 가능성을 확보하기 위해 표준화된 특징 추출 워크플로우가 필요하다는 점을 강조한다.

ABSTRACT

There is no consensus regarding the radiomic feature terminology, the underlying mathematics, or their implementation. This creates a scenario where features extracted using different toolboxes could not be used to build or validate the same model leading to a non-generalization of radiomic results. In this study, the image biomarker standardization initiative (IBSI) established phantom and benchmark values were used to compare the variation of the radiomic features while using 6 publicly available software programs and 1 in-house radiomics pipeline. All IBSI-standardized features (11 classes, 173 in total) were extracted. The relative differences between the extracted feature values from the different software and the IBSI benchmark values were calculated to measure the inter-software agreement. To better understand the variations, features are further grouped into 3 categories according to their properties: 1) morphology, 2) statistic/histogram and 3)texture features. While a good agreement was observed for a majority of radiomics features across the various programs, relatively poor agreement was observed for morphology features. Significant differences were also found in programs that use different gray level discretization approaches. Since these programs do not include all IBSI features, the level of quantitative assessment for each category was analyzed using Venn and the UpSet diagrams and also quantified using two ad hoc metrics. Morphology features earns lowest scores for both metrics, indicating that morphological features are not consistently evaluated among software programs. We conclude that radiomic features calculated using different software programs may not be identical and reliable. Further studies are needed to standardize the workflow of radiomic feature extraction.

연구 동기 및 목표

  • 표준화된 기준을 사용하여 소프트웨어 간 방사학적 특징 추출 일관성 평가.
  • 다양한 소프트웨어 툴킷과 내부 파이프라인 간 방사학적 특징의 변동성 원인 규명.
  • IBSI 표준화 조건 하에서 형태학적, 텍스처, 통계/히스토그램 특징의 신뢰성 평가.
  • 소프트웨어 간 다른 회색 수준 이산화 방법에 기인한 오차 양상 정량화.
  • 재현성과 모델 일반화를 향상시키기 위해 방사학적 워크플로우 표준화 기반 마련.

제안 방법

  • 모든 소프트웨어에서 동일한 IBSI 표준화된 페phantom을 사용하여 기준값을 제공하는 기준으로 사용.
  • 11개 클래스로 나누어 세 부류(형태학적, 통계/히스토그램, 텍스처 특징)로 분류된 총 173개의 IBSI 표준화 특징을 추출.
  • 소프트웨어에서 유도된 값과 IBSI 기준값 간 상대적 차이를 계산하여 소프트웨어 간 일치도 측정.
  • 특징 커버리지와 겹침을 분석하기 위해 Venn 및 UpSet 다이어그램 적용.
  • 각 부문별 특징 구현의 완전성과 일관성을 정량적으로 평가하기 위해 두 가지 임시 지표 도입.
  • 모든 도구에서 동일한 계산 파rameter(예: 보간, 이산화, 이웃 설정)를 표준화하여 알고리즘적 차이만 분리.

실험 결과

연구 질문

  • RQ1동일한 IBSI 표준화된 페phantom을 사용할 때, 다양한 소프트웨어 툴킷 간 방사학적 특징 값은 얼마나 일관성 있는가?
  • RQ2형태학적, 텍스처, 통계/히스토그램 특징 중 어느 부문에서 소프트웨어 간 일치도가 가장 높고 낮은가?
  • RQ3회색 수준 이산화 방법의 차이가 특징 값의 변동성에 어느 정도 기여하는가?
  • RQ4평가된 소프트웨어 도구들 간 IBSI 표준화 특징의 구현 완전도는 어느 정도인가?
  • RQ5재현성과 모델 일반화를 저해하는 방사학적 특징 추출의 주요 변동성 원인은 무엇인가?

주요 결과

  • 대부분의 방사학적 특징이 양호한 소프트웨어 간 일치도를 보였으며, 특히 텍스처 및 통계/히스토그램 부문에서 두드러졌다.
  • 형태학적 특징은 소프트웨어 간 일치도가 가장 낮았으며, 두 가지 임시 일관성 지표 모두에서 최저 점수를 기록했다.
  • 고정된 박스 크기(1.0)와 고정된 박스 수(6)를 사용하는 소프트웨어 간에 회색 수준 이산화에서 상당한 차이가 관찰되었다.
  • 형태학적 특징 중 27개(총 29개)만 평가된 도구들 간에 실제로 구현되었으며, 커버리지와 일관성 모두가 가장 낮았다.
  • LIFEx, CaPTk, A2가 가장 많은 IBSI 특징(총 173개)을 지원했으며, Pyradiomics와 SERA는 특정 특징 부문에서 지원이 제한적이었다.
  • GLCM, GLRLM, GLSZM 특징에서 동일한 알고리즘 기준을 사용하더라도 집계 방법과 이웃 정의의 차이로 인해 오차가 관찰되었다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.