[論文レビュー] Dataset Growth in Medical Image Analysis Research
本研究は、2011年から2018年のMICCAI会議論文を分析することで、MRI、CT、fMRI画像モodalitiesにおける被験者数のデータセットサイズの傾向を調査し、一貫した指数関数的成長が見られた。著者らは、7年間で中央値のデータセットサイズが3〜10倍に増加し、年間成長率が21〜31%であったと報告しており、これは査読プロセスが事実上、より大きなデータセットを要求する基準として機能していることを示唆している。
Medical image analysis studies usually require medical image datasets for training, testing and validation of algorithms. The need is underscored by the deep learning revolution and the dominance of machine learning in recent medical image analysis research. Nevertheless, due to ethical and legal constraints, commercial conflicts and the dependence on busy medical professionals, medical image analysis researchers have been described as "data starved". Due to the lack of objective criteria for sufficiency of dataset size, the research community implicitly sets ad-hoc standards by means of the peer review process. We hypothesize that peer review requires researchers to report the use of ever-increasing datasets as one condition for acceptance of their work to reputable publication venues. To test this hypothesis, we scanned the proceedings of the eminent MICCAI (Medical Image Computing and Computer-Assisted Intervention) conferences from 2011 to 2018. From a total of 2136 articles, we focused on 907 papers involving human datasets of MRI (Magnetic Resonance Imaging), CT (Computed Tomography) and fMRI (functional MRI) images. For each modality, for each of the years 2011-2018 we calculated the average, geometric mean and median number of human subjects used in that year's MICCAI articles. The results corroborate the dataset growth hypothesis. Specifically, the annual median dataset size in MICCAI articles has grown roughly 3-10 times from 2011 to 2018, depending on the imaging modality. Statistical analysis further supports the dataset growth hypothesis and reveals exponential growth of the geometric mean dataset size, with annual growth of about 21% for MRI, 24% for CT and 31% for fMRI. In slight analogy to Moore's law, the results can provide guidance about trends in the expectations of the medical image analysis community regarding dataset size.
研究の動機と目的
- 医学画像解析におけるデータセットサイズ要件が、査読基準の暗黙の規定によって時間経過とともに増加したかどうかを調査すること。
- 特に高影響力の国際会議を対象として、査読済み医学画像解析研究で使用されたデータセットサイズの成長傾向を定量化すること。
- 医学画像解析分野が、ムーアの法則に類似した暗黙の期待、すなわちより大きなデータセットを求める慣習を形成したかどうかを評価すること。
- 時間経過に伴うMRI、CT、fMRIの各画像モダリティにおけるデータセットサイズの成長パターンを分析すること。
- 倫理的・運用的制約が存在する中でも、医学画像解析研究におけるデータセットサイズの期待値が上昇しているという実証的証拠を提供すること。
提案手法
- 著者らは、MRI、CT、fMRIモダリティからの人間被験者データセットを用いた907件のMICCAI会議論文を収集・分析した。
- 各モダリティと年別に、関連論文の被験者数の平均値、幾何平均値、中央値を算出した。
- データセットサイズの成長傾向を評価するため、幾何平均値に指数関数的成長モデルをフィッティングする統計的分析を実施した。
- 本研究では、高影響力の医学画像解析研究の代表的サンプルとして、MICCAIのプロCEEDINGSを用いた。
- 臨床的・倫理的制約に関連する意義を保つために、本分析は人間被験者を対象とした研究に限定した。
- データセットサイズの成長率をモダリティごとに比較し、成長の違いを評価した。
実験結果
リサーチクエスチョン
- RQ12011年から2018年にかけて、査読済み研究で使用された医学画像データセットのサイズは著しく増加したか?
- RQ2MRI、CT、fMRIといった異なる画像モダリティは、データセットサイズの成長傾向に顕著な差異を示すか?
- RQ3医学画像解析研究において、データセットサイズが時間経過とともに指数関数的成長している証拠はあるか?
- RQ4査読プロセスが、より大きなデータセット要件を暗黙的に強いる程度はどの程度か?
- RQ5年次およびモダリティ別に、中央値、平均値、幾何平均値のデータセットサイズはどのように比較されるか?
主な発見
- 2011年から2018年にかけて、MICCAI論文における中央値のデータセットサイズは、画像モダリティに応じて3〜10倍に増加した。
- 幾何平均のデータセットサイズは、MRIで年間21%、CTで24%、fMRIで31%の割合で指数関数的に増加した。
- 統計的分析により、3つのモダリティすべてで明確で一貫した上行傾向が確認された。
- データ収集における継続的な倫理的・法的・運用的制約にもかかわらず、データセットサイズの増加が観察された。
- これらの結果から、査読プロセスが公式な基準がなくても、事実上より大きなデータセットを要求する基準を確立していることが示唆された。
- 観察された成長パターンは、指数関数的性質を示し、ムーアの法則に類似しており、データ要件の自己強化的トレンドを示している。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。