Skip to main content
QUICK REVIEW

[논문 리뷰] Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare Queries

Yiqiao Jin, Mohit Chandra|arXiv (Cornell University)|2023. 10. 19.
Artificial Intelligence in Healthcare and Education인용 수 10
한 줄 요약

논문은 XLingEval이라는 교차 언어 프레임워크와 XLingHealth라는 다국어 헬스케어 벤치마크를 도입하여 영어, 스페인어, 중국어, 힌디어에서 LLM을 평가하고, 정확성, 일관성, 검증 가능성에서 상당한 언어 간 차이를 드러냅니다.

ABSTRACT

Large language models (LLMs) are transforming the ways the general public accesses and consumes information. Their influence is particularly pronounced in pivotal sectors like healthcare, where lay individuals are increasingly appropriating LLMs as conversational agents for everyday queries. While LLMs demonstrate impressive language understanding and generation proficiencies, concerns regarding their safety remain paramount in these high-stake domains. Moreover, the development of LLMs is disproportionately focused on English. It remains unclear how these LLMs perform in the context of non-English languages, a gap that is critical for ensuring equity in the real-world use of these systems.This paper provides a framework to investigate the effectiveness of LLMs as multi-lingual dialogue systems for healthcare queries. Our empirically-derived framework XlingEval focuses on three fundamental criteria for evaluating LLM responses to naturalistic human-authored health-related questions: correctness, consistency, and verifiability. Through extensive experiments on four major global languages, including English, Spanish, Chinese, and Hindi, spanning three expert-annotated large health Q&A datasets, and through an amalgamation of algorithmic and human-evaluation strategies, we found a pronounced disparity in LLM responses across these languages, indicating a need for enhanced cross-lingual capabilities. We further propose XlingHealth, a cross-lingual benchmark for examining the multilingual capabilities of LLMs in the healthcare context. Our findings underscore the pressing need to bolster the cross-lingual capacities of these models, and to provide an equitable information ecosystem accessible to all.

연구 동기 및 목표

  • 고위험 분야에서 영어 이외의 LLM 평가를 통해 건강 정보에 대한 공정한 접근성을 촉진한다.
  • 정확성, 일관성, 검증 가능성에 초점을 맞춘 다국어 평가 프레임워크(XLingEval)를 제안한다.
  • 네 가지 널리 사용되는 언어에 걸친 다국어 헬스케어 벤치마크(XLingHealth)를 만든다.
  • 실세계 건강 QA 데이터셋에서 다수의 LLM에 대한 교차 언어 성능과 일반화를 평가한다.

제안 방법

  • 건강 질의에 대한 세 가지 핵심 평가 기준을 정의한다: 정확성, 일관성, 검증 가능성.
  • 다국어로 전문가-참고 진실과 대조하여 LLM 출력과를 비교하기 위해 자동화 및 인간 평가 구성요소를 갖춘 XLingEval을 개발한다.
  • 의료 전문가의 input과 함께 영어 건강 QA 데이터셋(HealthQA, LiveQA, MedicationQA)을 힌디어, 중국어, 스페인어로 번역하여 XLingHealth를 구축한다.
  • GPT-3.5와 MedAlpaca-30b를 사용한 다국어 실험을 수행하고 데이터셋과 언어 간 언어 차이를 분석한다.
  • 교차 언어 성능 차이의 통계적 유의성을 결정하기 위해 ANOVA, Tukey HSD, t-tests를 적용한다.
  • 표면적, 의미론적, 주제 수준의 일관성을 평가하기 위해 다수의 유사성 지표(n-gram, BERTScore, Sentence Embedding)와 주제 모델(LDA, HDP)을 사용한다.
  • 데이터셋 전반에서 모델을 올바른 주장과 잘못된 주장 구분자로 간주하여 검증 가능성을 평가한다.
Figure 1 . We present XLingEval , a comprehensive framework for assessing cross-lingual behaviors of LLMs for high risk domains such as healthcare. We present XLingHealth , a cross-lingual benchmark for healthcare queries.
Figure 1 . We present XLingEval , a comprehensive framework for assessing cross-lingual behaviors of LLMs for high risk domains such as healthcare. We present XLingHealth , a cross-lingual benchmark for healthcare queries.

실험 결과

연구 질문

  • RQ1LLM은 영어, 스페인어, 중국어, 힌디어에 걸친 건강 관련 질의에 대해 어떻게 성능을 보이나요?
  • RQ2정확성, 일관성, 검증 가능성은 건강 Q&A에서 교차 언어 차이를 보이나요?
  • RQ3XLingEval 프레임워크가 다국어 간 격차를 신뢰할 수 있게 탐지하고 교차 언어 건강 정보 접근성 개선을 안내할 수 있나요?
  • RQ4XLingHealth와 같은 다국어 벤치마크가 다른 도메인이나 모델에 일반화될 수 있나요?

주요 결과

  • 네 가지 언어에 걸쳐 정확성에 뚜렷한 차이가 있으며, GPT-3.5의 경우 비영어 질의가 영어보다 더 많은 잘못된 응답을 생성합니다.
  • 건강 데이터셋의 비영어 질의에서 GPT-3.5는 비영어 대비 영어보다 잘못된 답을 할 확률이 5.82배 더 높습니다.
  • 일관성 분석에서 힌디어는 최대 50.5%, 중국어는 영어와 비교해 특정 지표에서 28.3%의 성능 하락을 보입니다.
  • 검증 가능성은 중국어와 힌디어에서 현저히 약하며; 영어와 스페인어가 비교적 더 우수하게 나타납니다 (HealthQA: 영어 대 중국어/힌디어).
  • MedAlpaca-30b는 GPT-3.5와 다른 언어 차이 패턴을 보이며 모델 의존적 교차 언어 동작을 강조합니다.
  • ANOVA는 지표와 모델 전반에 걸쳐 통계적으로 유의한 언어 차이를 나타내며, 영어-스페인어가 종종 성능이 더 가깝고 다른 쌍은 더 큰 차이를 보입니다.
Figure 2 . Evaluation pipelines for correctness, consistency, and verifiability criteria in the XLingEval framework.
Figure 2 . Evaluation pipelines for correctness, consistency, and verifiability criteria in the XLingEval framework.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.