[論文レビュー] PMC-Patients: A Large-scale Dataset of Patient Summaries and Relations for Benchmarking Retrieval-based Clinical Decision Support Systems
本稿では、PubMed Centralの症例報告から抽出した167,000件の患者要約、310万件の患者-論文関連性アノテーション、293,000件の患者-患者類似性アノテーションを備えた大規模かつ公開可能なデータセット、PMC-Patientsを紹介する。2つのベンチマークタスク、Patient-to-Article Retrieval (ReCDS-PAR) および Patient-to-Patient Retrieval (ReCDS-PPR) を確立し、既存のトップレベルベースラインでさえReCDS-PARではP@10が16%、ReCDS-PPRではP@10が6%にとどまることを示し、臨床意思決定支援におけるより優れた検索手法の必要性を浮き彫りにしている。
Objective: Retrieval-based Clinical Decision Support (ReCDS) can aid clinical workflow by providing relevant literature and similar patients for a given patient. However, the development of ReCDS systems has been severely obstructed by the lack of diverse patient collections and publicly available large-scale patient-level annotation datasets. In this paper, we aim to define and benchmark two ReCDS tasks: Patient-to-Article Retrieval (ReCDS-PAR) and Patient-to-Patient Retrieval (ReCDS-PPR) using a novel dataset called PMC-Patients. Methods: We extract patient summaries from PubMed Central articles using simple heuristics and utilize the PubMed citation graph to define patient-article relevance and patient-patient similarity. We also implement and evaluate several ReCDS systems on the PMC-Patients benchmarks, including sparse retrievers, dense retrievers, and nearest neighbor retrievers. We conduct several case studies to show the clinical utility of PMC-Patients. Results: PMC-Patients contains 167k patient summaries with 3.1M patient-article relevance annotations and 293k patient-patient similarity annotations, which is the largest-scale resource for ReCDS and also one of the largest patient collections. Human evaluation and analysis show that PMC-Patients is a diverse dataset with high-quality annotations. The evaluation of various ReCDS systems shows that the PMC-Patients benchmark is challenging and calls for further research. Conclusion: We present PMC-Patients, a large-scale, diverse, and publicly available patient summary dataset with the largest-scale patient-level relation annotations. Based on PMC-Patients, we formally define two benchmark tasks for ReCDS systems and evaluate various existing retrieval methods. PMC-Patients can largely facilitate methodology research on ReCDS systems and shows real-world clinical utility.
研究の動機と目的
- 検索ベースの臨床意思決定支援(ReCDS)システムの学習および評価に用いることのできる、多様性に富み、大規模かつ公開可能な患者レベルのデータセットの不足に対処すること。
- ReCDSシステムの評価を目的とした2つの新しいベンチマークタスク、Patient-to-Article Retrieval (ReCDS-PAR) および Patient-to-Patient Retrieval (ReCDS-PPR) を定義および確立すること。
- 研究手法の開発およびReCDSにおける臨床的有用性の評価を支援する、高品質で多様性に富み、スケーラブルなデータセットを提供すること。
- すべてのデータおよびコードをhttps://github.com/pmc-patients/pmc-patientsに公開し、臨床NLPおよび情報検索分野における研究開発を加速すること。
提案手法
- 患者要約は、PubMed Centralの症例報告から、患者記述を特定するための単純なヒューリスティクスを用いて抽出された。
- 患者-論文関連性は、PubMedの引用グラフを用いて定義され、患者症例と論文の間で共通の引用が存在する場合、関連性があるとみなされた。
- 患者-患者類似性は、同じ論文に対して共通の引用が存在するという事実に基づき推定され、共引用が臨床的類似性を示すと仮定した。
- スパars、ディンス、および近隣検索モデルを実装し、2つのベンチマークタスクで評価することで、ベースライン性能を確立した。
- データセットのアノテーションの質と多様性を検証するために、人間による評価を実施した。
- データセットは、https://github.com/pmc-patients/pmc-patients に完全なコードとデータとともに公開され、コミュニティ研究を支援する。
実験結果
リサーチクエスチョン
- RQ1既存の検索モデルは、特定の患者要約に対して関連する臨床論文をどれほど効果的に特定できるか。
- RQ2医学文献における共通の引用から、患者-患者類似性をどれほど信頼性を持って推定できるか。
- RQ3大規模かつ現実世界の臨床的患者検索タスクにおいて、現在の検索手法の性能の上限はどの程度か。
- RQ4患者の疾患状態や医学専門分野の観点から、PMC-Patientsデータセットはどれほど多様で臨床的に代表的か。
- RQ5提案されたベンチマークは、現実世界の臨床意思決定支援シナリオの課題を効果的に捉えているか。
主な発見
- PMC-Patientsには167,000件の患者要約、310万件の患者-論文関連性アノテーション、293,000件の患者-患者類似性アノテーションが含まれており、同種のデータセットの中で最大規模である。
- 人間による評価により、データセットが多様性に富み、臨床ベンチマークに適した高品質なアノテーションを含んでいることが確認された。
- 最良のベースライン検索モデルは、ReCDS-PARタスクでP@10が16%、R@1kが63%にとどまり、改善の余地が著しくあることが示された。
- ReCDS-PPRタスクでは、最良のベースラインがP@10が6%、R@1kが80%を記録し、患者-患者検索が依然として困難な課題であることが示された。
- 事例研究により、関連する文献や類似患者症例を活用することで、実際の診断および治療意思決定を支援するという、データセットの臨床的有用性が示された。
- データセットおよびコードの公開により、再現可能な研究が可能となり、検索ベースの臨床意思決定支援分野におけるイノベーションが加速された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。