Skip to main content
QUICK REVIEW

[論文レビュー] A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor?

Yunfei Xie, Jianzhen Wu|arXiv (Cornell University)|Sep 23, 2024
Artificial Intelligence in Healthcare and Education被引用数 14
ひとこと要約

この論文は OpenAI の o1 モデルを 37 の医療データセットで、理解、推論、 multilinguality において評価し、医療理解と推論の改善を示す一方で、幻覚、 多言語の課題、およびメトリクスの不整合を指摘している。

ABSTRACT

Large language models (LLMs) have exhibited remarkable capabilities across various domains and tasks, pushing the boundaries of our knowledge in learning and cognition. The latest model, OpenAI's o1, stands out as the first LLM with an internalized chain-of-thought technique using reinforcement learning strategies. While it has demonstrated surprisingly strong capabilities on various general language tasks, its performance in specialized fields such as medicine remains unknown. To this end, this report provides a comprehensive exploration of o1 on different medical scenarios, examining 3 key aspects: understanding, reasoning, and multilinguality. Specifically, our evaluation encompasses 6 tasks using data from 37 medical datasets, including two newly constructed and more challenging question-answering (QA) tasks based on professional medical quizzes from the New England Journal of Medicine (NEJM) and The Lancet. These datasets offer greater clinical relevance compared to standard medical QA benchmarks such as MedQA, translating more effectively into real-world clinical utility. Our analysis of o1 suggests that the enhanced reasoning ability of LLMs may (significantly) benefit their capability to understand various medical instructions and reason through complex clinical scenarios. Notably, o1 surpasses the previous GPT-4 in accuracy by an average of 6.2% and 6.6% across 19 datasets and two newly created complex QA scenarios. But meanwhile, we identify several weaknesses in both the model capability and the existing evaluation protocols, including hallucination, inconsistent multilingual ability, and discrepant metrics for evaluation. We release our raw data and model outputs at https://ucsc-vlaa.github.io/o1_medicine/ for future research.

研究の動機と目的

  • o1 モデルの高度な推論が医療領域へ転送されるかを評価する。
  • 多様なデータセットを用いて医療理解、推論、及び多言語能力を評価する。
  • 複数の医療タスクを通じて o1 を GPT-4、GPT-3.5、オープンソースのベースラインと比較する。
  • 医療文脈におけるモデルの弱点と現行評価プロトコルを特定し、臨床AI開発の未来を導く。

提案手法

  • 3つの医療側面を横断する37データセット(35既存+2新規)の広範な評価スイートを組み立てる。
  • 直接 prompting、チェーン・オブ・ソウト(CoT)、Few-shot prompting の3つの prompting 戦略を使用し、o1 の内部 CoT 訓練を踏まえた CoT の影響を評価する。
  • o1 を GPT-4、GPT-3.5、MEDITRON-70B、Llama3-8B と、6つのタスクと3つの側面にわたって比較する。
  • 精度、F1、BLEU、ROUGE、AlignScore、Mauve などの指標を用いて、理解、推論、多言語性といった異なるタスクタイプを評価する。
  • CoT、Self-Consistency、Reflex など追加 prompting で結果を分析し、 prompting の影響を研究する。

実験結果

リサーチクエスチョン

  • RQ1o1 の内部のチェーン・オブ・ソウトと強化学習訓練は、従来のモデルと比較して臨床理解と推論を改善できるか。
  • RQ2o1 は 医療理解、推論、多言語タスクにおいて、GPT-4、GPT-3.5、オープンソースのベースラインと比較してどう性能か。
  • RQ3医療文脈における o1 の限界(幻覚、多言語の課題など)と、評価指標がモデルランキングに与える影響は何か。

主な発見

  • o1 は多くのデータセットで理解と特定の推論タスクで GPT-4 および GPT-3.5 を上回る傾向。
  • 新規 NEJMQA および LancetQA タスクで o1 は GPT-4 および GPT-3.5 より著しい正確性の向上を示す。
  • o1 は自由形式生成タスクで ROUGE-1 スコアが高く、要約品質が改善される。
  • 幻覚は依然課題で、複雑な多言語状況は多言語推論のギャップを露呈する。
  • 評価指標はモデル間で一貫したランキングを与えず、堅牢でドメイン特化の指標の必要性を浮き彫りにする。
  • CoT prompting は o1 の医療知識タスクを改善する可能性があるが、全タスクタイプで普遍的ではない。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。