[論文レビュー] Benchmarking features from different radiomics toolkits / toolboxes using Image Biomarkers Standardization Initiative
本研究では、既知のベンチマーク値を持つデジタルフォントを用いて、6つの公開ツールキットおよび1つの自社開発パイプラインを対象に、173のIBSI標準化済みラジオミクス特徴量をベンチマーク化した。結果として、特に形状特徴量およびグレーレベルの離散化手法において、ソフトウェア間で顕著な差異が認められ、再現性およびラジオミクス研究における一般化可能性を確保するための標準化された特徴抽出ワークフローの必要性が浮き彫りになった。
There is no consensus regarding the radiomic feature terminology, the underlying mathematics, or their implementation. This creates a scenario where features extracted using different toolboxes could not be used to build or validate the same model leading to a non-generalization of radiomic results. In this study, the image biomarker standardization initiative (IBSI) established phantom and benchmark values were used to compare the variation of the radiomic features while using 6 publicly available software programs and 1 in-house radiomics pipeline. All IBSI-standardized features (11 classes, 173 in total) were extracted. The relative differences between the extracted feature values from the different software and the IBSI benchmark values were calculated to measure the inter-software agreement. To better understand the variations, features are further grouped into 3 categories according to their properties: 1) morphology, 2) statistic/histogram and 3)texture features. While a good agreement was observed for a majority of radiomics features across the various programs, relatively poor agreement was observed for morphology features. Significant differences were also found in programs that use different gray level discretization approaches. Since these programs do not include all IBSI features, the level of quantitative assessment for each category was analyzed using Venn and the UpSet diagrams and also quantified using two ad hoc metrics. Morphology features earns lowest scores for both metrics, indicating that morphological features are not consistently evaluated among software programs. We conclude that radiomic features calculated using different software programs may not be identical and reliable. Further studies are needed to standardize the workflow of radiomic feature extraction.
研究の動機と目的
- 標準化されたベンチマークを用いて、ソフトウェア間の一貫性を評価すること。
- 異なるソフトウェアツールキットおよび自社開発パイプライン間でのラジオミクス特徴量のばらつきの原因を同定すること。
- IBSI標準化のもとでの形状、テクスチャ、統計/ヒストグラム特徴量の信頼性を評価すること。
- 異なるグレーレベルの離散化アプローチに起因する差異を定量化すること。
- 再現性およびモデルの一般化可能性を向上させるために、ラジオミクスワークフローの標準化の基盤を提供すること。
提案手法
- ラジオミクス特徴量の計算のベンチマークとして、既知の真値を有するIBSIデジタルフォントを用いた。
- 11のクラスに分類された3つのカテゴリ(形状、統計/ヒストグラム、テクスチャ)に分け、全173のIBSI標準化特徴量を抽出した。
- ソフトウェアが算出した値とIBSIベンチマーク値との相対差を計算し、ソフトウェア間の一致度を測定した。
- Venn図およびUpSet図を用いて、ソフトウェアツール間での特徴量のカバー範囲および重複度を分析した。
- 各カテゴリごとの特徴量実装の完全性および一貫性を定量的に評価するための2つのアドホック指標を導入した。
- 補間、離散化、近隣設定などの計算パラメータをすべてのツールで標準化し、アルゴリズム的差異を隔離した。
実験結果
リサーチクエスチョン
- RQ1同じIBSI標準化フォントを用いた場合、異なるソフトウェアツールキット間でラジオミクス特徴量の値はどの程度一貫しているか?
- RQ2形状、テクスチャ、統計/ヒストグラムの各カテゴリの中で、ソフトウェア間一致度が最も高く・低かったのはどれか?
- RQ3グレーレベルの離散化手法の違いが、特徴量の値のばらつきにどの程度寄与しているか?
- RQ4評価対象のソフトウェアツールでは、IBSI標準化特徴量の実装はどの程度完全であったか?
- RQ5再現性およびモデルの一般化を阻害するラジオミクス特徴量抽出における主なばらつき要因は何か?
主な発見
- 大部分のラジオミクス特徴量で、特にテクスチャおよび統計/ヒストグラムカテゴリにおいて、良好なソフトウェア間一貫性が確認された。
- 形状特徴量はソフトウェア間で最も低い一致度を示し、2つのアドホック一貫性指標でも最低得点を記録した。
- グレーレベルの離散化に固定ビンサイズ(1.0)と固定ビン数(6)を用いたソフトウェア間で顕著な差が認められた。
- 形状特徴量のうち27/29が評価対象ツールで実装されており、カバー範囲および一貫性が最低水準であった。
- LIFEx、CaPTk、A2が最も多くのIBSI特徴量(合計173)を実装したが、PyradiomicsおよびSERAは特定の特徴量カテゴリでサポートが限定的であった。
- 同じアルゴリズム的リファレンスを用いても、GLCM、GLRLM、GLSZM特徴量に差異が認められ、これは集約方法や近隣定義の違いに起因した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。