Skip to main content
QUICK REVIEW

[논문 리뷰] Stochastic Parrots or ICU Experts? Large Language Models in Critical Care Medicine: A Scoping Review

Tongyue Shi, Jun Ma|arXiv (Cornell University)|2024. 07. 27.
Machine Learning in Healthcare인용 수 4
한 줄 요약

이 범위 검토는 2019–2024년 사이의 24篇의 연구를 분석하여 대규모 언어 모델(Large Language Models, LLMs)이 중환자 집중 치료 의학(Critical Care Medicine, CCM) 분야에 어떻게 적용되고 있는지 조사한다. LLMs는 임상 의사결정 지원, 문서화, 의료 교육 분야에서 잠재력을 보이고 있으나, 환상 생성, 낮은 해석 가능성, 편향, 윤리적 리스크 등의 과제에 직면해 있으며, 이에 따라 신뢰성 향상, 지식 통합, 윤리적 프레임워크 개선이 필요하다고 제안한다.

ABSTRACT

With the rapid development of artificial intelligence (AI), large language models (LLMs) have shown strong capabilities in natural language understanding, reasoning, and generation, attracting amounts of research interest in applying LLMs to health and medicine. Critical care medicine (CCM) provides diagnosis and treatment for critically ill patients who often require intensive monitoring and interventions in intensive care units (ICUs). Can LLMs be applied to CCM? Are LLMs just like stochastic parrots or ICU experts in assisting clinical decision-making? This scoping review aims to provide a panoramic portrait of the application of LLMs in CCM. Literature in seven databases, including PubMed, Embase, Scopus, Web of Science, CINAHL, IEEE Xplore, and ACM Digital Library, were searched from January 1, 2019, to June 10, 2024. Peer-reviewed journal and conference articles that discussed the application of LLMs in critical care settings were included. From an initial 619 articles, 24 were selected for final review. This review grouped applications of LLMs in CCM into three categories: clinical decision support, medical documentation and reporting, and medical education and doctor-patient communication. LLMs have advantages in handling unstructured data and do not require manual feature engineering. Meanwhile, applying LLMs to CCM faces challenges, including hallucinations, poor interpretability, bias and alignment challenges, and privacy and ethics issues. Future research should enhance model reliability and interpretability, integrate up-to-date medical knowledge, and strengthen privacy and ethical guidelines. As LLMs evolve, they could become key tools in CCM to help improve patient outcomes and optimize healthcare delivery. This study is the first review of LLMs in CCM, aiding researchers, clinicians, and policymakers to understand the current status and future potentials of LLMs in CCM.

연구 동기 및 목표

  • 대규모 언어 모델(Large Language Models, LLMs)이 중환자 집중 치료 의학(Critical Care Medicine, CCM) 분야에 어떻게 적용되고 있는지 현재의 전반적 풍경을 맵핑하기 위해.
  • 임상 의사결정 지원, 의료 문서화, 의료 교육 등 LLMs가 적용되고 있는 주요 분야를 특정하기 위해.
  • 고위험도의 ICU 환경에서 LLMs를 도입할 경우 발생하는 이점과 리스크를 평가하기 위해.
  • 환상 생성, 해석 불가능성 부족, 편향, 윤리적 우려 등 임상 분야에서 LLM 사용 시 주요 과제를 부각하기 위해.
  • 미래 연구를 이끌기 위해 모델의 신뢰성, 실시간 지식 통합, 철저한 개인정보 보호 및 윤리 기준을 우선순위로 삼기 위해.

제안 방법

  • PubMed, Embase, Scopus, Web of Science, CINAHL, IEEE Xplore, ACM 디지털 라이브러리 등 7개의 데이터베이스를 대상으로 체계적 검색을 수행하였다.
  • 포괄 기준: 2019년 1월 1일부터 2024년 6월 10일까지 출판된 동료 심사 기반의 학술지 및 컫퍼런스 논문으로, CCM 분야에서 LLM의 적용에 중점을 두었다.
  • 제목/초록 및 전문 검토를 통해 619篇의 논문을 선별하였으며, 최종 분석 대상으로 24篇의 연구를 확정하였다.
  • LLM 응용 분야를 세 가지 영역으로 분류: 임상 의사결정 지원, 의료 문서화 및 보고서 작성, 의료 교육 및 의사-환자 소통.
  • 주제별로 통합된 결과를 분석하여, ICU 내 LLM 도입 시 기술적, 윤리적, 임상적 과제를 강조하였다.
  • 신뢰성, 해석 가능성, 윤리적 거버넌스 분야의 격차를 기반으로 향후 연구를 위한 권고안을 제시하였다.

실험 결과

연구 질문

  • RQ1최근 문헌에서 보고된 대규모 언어 모델(Large Language Models, LLMs)의 주요 응용 분야는 무엇인가?
  • RQ2대규모 언어 모델(Large Language Models, LLMs)은 중환자 치료 환경 내 비정형 임상 데이터를 어떻게 처리하는가?
  • RQ3ICU 환경에서 LLMs를 도입함에 있어 관련된 주요 기술적 및 윤리적 과제는 무엇인가?
  • RQ4LLMs는 중환자 치료 분야에서 임상 의사결정 지원, 문서화 효율성 향상, 또는 의료 교육 향상에 어느 정도 기여하는가?
  • RQ5LLMs의 안전성, 신뢰성, 윤리적 사용을 향상시키기 위해 향후 어떤 연구 방향이 필요할까?

주요 결과

  • LLMs는 CCM 응용 분야에서 비정형 임상 데이터를 처리하는 데 강력한 능력을 보이며, 수동적인 특징 엔지니어링의 필요성을 줄였다.
  • CCM 분야에서 LLM 응용의 대부분은 세 가지 영역으로 분류되었다: 임상 의사결정 지원(예: 진단 및 치료 제안), 의료 문서화(예: 자동 노트 생성), 의료 교육(예: 훈련 및 소통 도구).
  • 환상 생성—사실과 맞지 않거나 허구적인 임상 정보를 생성하는 것—은 진단 및 치료 제안 작업에서 특히 두드러진 위험으로 여러 연구에서 확인되었다.
  • 모델의 낮은 해석 가능성과 추론 과정의 투명성 부족은 임상적 신뢰도 향상과 도입을 저해하는 데 지속적으로 언급된 장벽이었다.
  • 학습 데이터의 편향과 임상 지침과의 불일치는 출력의 공정성과 안전성에 영향을 주는 주요 위험 요소로 규명되었다.
  • 환자 데이터 유출 및 규제 감시 부족 등 개인정보 및 윤리적 우려는 실생활 도입을 저해하는 주요 장벽으로 자주 제기되었다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.