Skip to main content
QUICK REVIEW

[논문 리뷰] Evaluation of GPT-3.5 and GPT-4 for supporting real-world information needs in healthcare delivery

Debadutta Dash, Rahul Thapa|arXiv (Cornell University)|2023. 04. 26.
Artificial Intelligence in Healthcare and Education인용 수 21
한 줄 요약

연구는 의료 현장에서 의사의 질문에 대한 정보학 상담 보조로 GPT-3.5와 GPT-4를 평가하고, 전문가 보고서와의 합치성은 제한적이며 다수의 해롭다는 신호는 없었고, Prompt 엔지니어링 및 모델 맞춤화의 필요성을 강조한다.

ABSTRACT

Despite growing interest in using large language models (LLMs) in healthcare, current explorations do not assess the real-world utility and safety of LLMs in clinical settings. Our objective was to determine whether two LLMs can serve information needs submitted by physicians as questions to an informatics consultation service in a safe and concordant manner. Sixty six questions from an informatics consult service were submitted to GPT-3.5 and GPT-4 via simple prompts. 12 physicians assessed the LLM responses' possibility of patient harm and concordance with existing reports from an informatics consultation service. Physician assessments were summarized based on majority vote. For no questions did a majority of physicians deem either LLM response as harmful. For GPT-3.5, responses to 8 questions were concordant with the informatics consult report, 20 discordant, and 9 were unable to be assessed. There were 29 responses with no majority on "Agree", "Disagree", and "Unable to assess". For GPT-4, responses to 13 questions were concordant, 15 discordant, and 3 were unable to be assessed. There were 35 responses with no majority. Responses from both LLMs were largely devoid of overt harm, but less than 20% of the responses agreed with an answer from an informatics consultation service, responses contained hallucinated references, and physicians were divided on what constitutes harm. These results suggest that while general purpose LLMs are able to provide safe and credible responses, they often do not meet the specific information need of a given question. A definitive evaluation of the usefulness of LLMs in healthcare settings will likely require additional research on prompt engineering, calibration, and custom-tailoring of general purpose models.

연구 동기 및 목표

  • 두 대형 언어 모델(GPT-3.5와 GPT-4)이 의사가 제출한 정보학 질문에 안전하게 대답할 수 있는지 여부를 평가한다.
  • LLM 응답과 확립된 정보학 상담 보고서 사이의 합치성을 평가한다.
  • 실제 임상 질의에서 환자 해를 초래할 수 있는 위험 및 환각 등 안전 문제를 식별한다.

제안 방법

  • 간단한 프롬프트로 정보학 상담 서비스의 66개 의사 질문을 GPT-3.5 및 GPT-4에 제출한다.
  • 12명의 의사가 LLM 응답이 환자 해를 초래하는지, 그리고 정보학 상담 보고서와 합치하는지 평가하도록 한다.
  • 안전성과 합치성을 판단하기 위해 의사 평가를 다수결로 요약한다.
  • 각 모델별로 합치/불일치/평가 불가 여부의 수를 보고한다.

실험 결과

연구 질문

  • RQ1GPT-3.5와 GPT-4가 의료 제공의 실제 의사 정보 필요에 안전하게 응답할 수 있는가?
  • RQ2LLM 응답이 확립된 정보학 상담 보고서와 어느 정도 합치하는가?
  • RQ3임상 질의를 위한 LLM 출력에서 관찰되는 해롭다, 환각, 또는 불일치의 패턴은 무엇인가?

주요 결과

  • 어떤 의사도 어느 LLM 응답이 해롭다고 다수의 의견을 낸 경우가 없었다.
  • GPT-3.5: 합치 8, 불일치 20, 평가 불가 9; 합의/비합의/평가 불가 다수결이 아닌 경우가 29.
  • GPT-4: 합치 13, 불일치 15, 평가 불가 3; 합의/비합의/평가 불가 다수결이 아닌 경우가 35.
  • 두 LLM의 응답은 명백한 해롭다기보다는 환각적 참조를 포함하는 경우가 많았고 정보학 상담 보고서와 일치하지 않는 경우가 많았다.
  • 정보학 상담 서비스의 답변에 동의한 응답은 20% 미만이었다.
  • 일반-purpose LLM은 안전하나 추가 프롬프트 엔지니어링과 맞춤화 없이는 신뢰할 수 있을 만큼 유용하지 않음을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.