[論文レビュー] Fine-tuning Large Language Model (LLM) Artificial Intelligence Chatbots in Ophthalmology and LLM-based evaluation using GPT-4
本稿は眼科の質問に対して複数のLLMチャットボットをファインチューニングし、臨床医のランキングと照合したGPT-4ベースのスコアリングを評価しており、高い一致を示す一方、いくつかのモデルに臨床的な不正確さを指摘している。
Purpose: To assess the alignment of GPT-4-based evaluation to human clinician experts, for the evaluation of responses to ophthalmology-related patient queries generated by fine-tuned LLM chatbots. Methods: 400 ophthalmology questions and paired answers were created by ophthalmologists to represent commonly asked patient questions, divided into fine-tuning (368; 92%), and testing (40; 8%). We find-tuned 5 different LLMs, including LLAMA2-7b, LLAMA2-7b-Chat, LLAMA2-13b, and LLAMA2-13b-Chat. For the testing dataset, additional 8 glaucoma QnA pairs were included. 200 responses to the testing dataset were generated by 5 fine-tuned LLMs for evaluation. A customized clinical evaluation rubric was used to guide GPT-4 evaluation, grounded on clinical accuracy, relevance, patient safety, and ease of understanding. GPT-4 evaluation was then compared against ranking by 5 clinicians for clinical alignment. Results: Among all fine-tuned LLMs, GPT-3.5 scored the highest (87.1%), followed by LLAMA2-13b (80.9%), LLAMA2-13b-chat (75.5%), LLAMA2-7b-Chat (70%) and LLAMA2-7b (68.8%) based on the GPT-4 evaluation. GPT-4 evaluation demonstrated significant agreement with human clinician rankings, with Spearman and Kendall Tau correlation coefficients of 0.90 and 0.80 respectively; while correlation based on Cohen Kappa was more modest at 0.50. Notably, qualitative analysis and the glaucoma sub-analysis revealed clinical inaccuracies in the LLM-generated responses, which were appropriately identified by the GPT-4 evaluation. Conclusion: The notable clinical alignment of GPT-4 evaluation highlighted its potential to streamline the clinical evaluation of LLM chatbot responses to healthcare-related queries. By complementing the existing clinician-dependent manual grading, this efficient and automated evaluation could assist the validation of future developments in LLM applications for healthcare.
研究の動機と目的
- ファインチューニングされたLLMチャットボットによって生成された眼科のQ&Aに対する、GPT-4ベースの評価と人間の臨床医専門家との整合性を評価する。
- 眼科の患者質問を対象に複数のLLMをファインチューニングしてベンチマークデータセットを作成する。
- 正確さ、関連性、安全性、理解しやすさを含むGPT-4主導の臨床評価基準を用いて応答を評価する。
提案手法
- 眼科専門医が代表的な患者の問い合わせを反映した400の眼科質問と対応する回答を作成する。
- ファインチューニング用(368問)とテスト用(40問)のセットに分割し、テストには緑内障のQ&Aペアを8組含める。
- 5つのLLMをファインチューニングする:LLAMA2-7b、LLAMA2-7b-Chat、LLAMA2-13b、LLAMA2-13b-Chat(および変種)。
- 5つのファインチューニング済みLLMからテストデータセットへの200件の応答を生成し、評価に供する。
- 臨床的正確性、関連性、患者の安全性、理解の容易さに焦点を当てたカスタマイズされた臨床評価基準を用いてGPT-4評価を導く。
- 臨床整合性を評価するため、GPT-4評価結果を5名の臨床医のランキングと比較する。
実験結果
リサーチクエスチョン
- RQ1ファインチューニング済みLLMによる眼科のQ&Aを評価する際に、GPT-4評価は人間の臨床医のランキングとどれくらい一致するか。
- RQ2眼科の質問応答において、どのLLM構成が最も良い臨床的整合性と性能を提供するか。
- RQ3臨床的にLLMsがどこで不足するか(例:緑内障)と、GPT-4が不正確さをどれだけ効果的に特定できるか。
主な発見
- GPT-4評価は人間の臨床医のランキングと高い一致を示し、Spearman 0.90と Kendall Tau 0.80である。
- Cohen Kappa agreement between GPT-4 evaluation and clinician rankings was more modest at 0.50.
- GPT-4評価は緑内障を含むLLM生成応答の臨床的不正確さを特定した。
- 全てのファインチューニングモデルの中で、GPT-3.5がGPT-4評価で最高の87.1%、次いでLLAMA2-13bが80.9%、LLAMA2-13b-chatが75.5%、LLAMA2-7b-Chatが70%、LLAMA2-7bが68.8%である。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。