[論文レビュー] A Comparative Study of Open-Source Large Language Models, GPT-4 and Claude 2: Multiple-Choice Test Taking in Nephrology
本論文は、腎臓病学のMCQにおいてオープンソースのLLMをGPT-4およびClaude 2と比較し、GPT-4とClaude 2がオープンソースモデルを大きく上回ることを示している。
In recent years, there have been significant breakthroughs in the field of natural language processing, particularly with the development of large language models (LLMs). These LLMs have showcased remarkable capabilities on various benchmarks. In the healthcare field, the exact role LLMs and other future AI models will play remains unclear. There is a potential for these models in the future to be used as part of adaptive physician training, medical co-pilot applications, and digital patient interaction scenarios. The ability of AI models to participate in medical training and patient care will depend in part on their mastery of the knowledge content of specific medical fields. This study investigated the medical knowledge capability of LLMs, specifically in the context of internal medicine subspecialty multiple-choice test-taking ability. We compared the performance of several open-source LLMs (Koala 7B, Falcon 7B, Stable-Vicuna 13B, and Orca Mini 13B), to GPT-4 and Claude 2 on multiple-choice questions in the field of Nephrology. Nephrology was chosen as an example of a particularly conceptually complex subspecialty field within internal medicine. The study was conducted to evaluate the ability of LLM models to provide correct answers to nephSAP (Nephrology Self-Assessment Program) multiple-choice questions. The overall success of open-sourced LLMs in answering the 858 nephSAP multiple-choice questions correctly was 17.1% - 25.5%. In contrast, Claude 2 answered 54.4% of the questions correctly, whereas GPT-4 achieved a score of 73.3%. We show that current widely used open-sourced LLMs do poorly in their ability for zero-shot reasoning when compared to GPT-4 and Claude 2. The findings of this study potentially have significant implications for the future of subspecialty medical training and patient care.
研究の動機と目的
- nephSAP MCQ におけるオープンソース LLM の腎臓病学知識能力を評価する。
- 腎臓病学データセットに対するゼロショット性能で、オープンソース LLM を GPT-4 および Claude 2 と比較する。
- モデルの説明の質と真値との意味的整合性を分析する。
提案手法
- Koala 7B、Falcon 7B、Stable-Vicuna 13B、Orca Mini 13B を GPT-4 および Claude 2 に対して、858 の nephSAP MCQ で評価する。
- Context、Question、Choices を連結してオープンソース LLM のフォワードパス用の入力プロンプトを構築する。
- 正解予測を決定するために正規表現ベースの抽出を用いて自動出力を解析し、予測された回答を基準真値と比較する。
- 正答数として生の正答率を測定し、腎臓病学のサブトピックごとにトピック別の性能を算出する。
- BLEU、WER、およびコサイン類似度を用いて、基準真の説明と比較して説明を評価する。)
実験結果
リサーチクエスチョン
- RQ1open-source LLM は nephSAP 腎臟病学 MCQ において GPT-4 および Claude 2 と比較してどのように性能を示すか?
- RQ2腎臓病学におけるオープンソース LLM のトピック別の長所と弱点は何か?
- RQ3基準真値との比較における説明の質はオープンソース LLM でどの程度か?
- RQ4オープンソース LLM と GPT-4/Claude 2 との間の性能差を説明しうる要因は何か?
主な発見
- GPT-4 は 629 問の正解(73.3%)、Claude 2 は 467 問(54.4%)、オープンソース LLM は 17.1% から 25.5% の正解率の範囲だった。
- オープンソースモデルの中では、Vicuna が最高スコアで 219 問正解(25.5%)、Koala が続き 204 問正解(23.8%)。
- Falcon が 155 問正解(18.1%)、Orca-Mini が 147 問正解(17.1%)、Koala が 204 問正解(23.8%)。
- 問題構造を考慮するとランダム推測は 23.8% を生むため、オープンソースモデルの中ではKoala のみがランダム期待をわずかに上回った。
- オープンソースモデルの説明は低いBLEUスコアを示し(例:Vicuna 9%、Falcon 8%、Orca-Mini 5%、Koala 5%)、コサイン類似度スコアもモデル間で最適とは言えなかった。
- GPT-4 と Claude 2 は全ての腎臓病学トピックでオープンソースモデルを上回り、GPT-4 はほとんどのトピックでほぼ人間の性能を示した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。