Skip to main content
QUICK REVIEW

[논문 리뷰] Evaluating LLMs on Document-Based QA: Exact Answer Selection and Numerical Extraction using Cogtale dataset

Zafaryab Rasool, Stefanus Kurniawan|arXiv (Cornell University)|2023. 11. 14.
Topic Modeling인용 수 7
한 줄 요약

이 논문은 CogTale 데이터셋을 사용하여 문서 기반 질의응답 작업에서 GPT-4와 GPT-3.5의 성능을 평가한다. 특히 정확한 답변 선택(단일 선택, 예/아니요, 다중 선택)과 수치 추출을 중심으로 제로샷 설정에서 평가한다. GPT-4는 단일 선택 및 예/아니요 질문에서 뛰어난 성능을 보였지만, 다중 선택 및 수치 추출 작업에서 성능 저하가 심각하게 나타나, 의료 메타분석과 같은 핵심 응용 분야에서 정확한 정보 검색 능력에 한계가 있음을 시사한다.

ABSTRACT

Document-based Question-Answering (QA) tasks are crucial for precise information retrieval. While some existing work focus on evaluating large language models performance on retrieving and answering questions from documents, assessing the LLMs performance on QA types that require exact answer selection from predefined options and numerical extraction is yet to be fully assessed. In this paper, we specifically focus on this underexplored context and conduct empirical analysis of LLMs (GPT-4 and GPT-3.5) on question types, including single-choice, yes-no, multiple-choice, and number extraction questions from documents in zero-shot setting. We use the CogTale dataset for evaluation, which provide human expert-tagged responses, offering a robust benchmark for precision and factual grounding. We found that LLMs, particularly GPT-4, can precisely answer many single-choice and yes-no questions given relevant context, demonstrating their efficacy in information retrieval tasks. However, their performance diminishes when confronted with multiple-choice and number extraction formats, lowering the overall performance of the model on this task, indicating that these models may not yet be sufficiently reliable for the task. This limits the applications of LLMs on applications demanding precise information extraction from documents, such as meta-analysis tasks. These findings hinge on the assumption that the retrievers furnish pertinent context necessary for accurate responses, emphasizing the need for further research. Our work offers a framework for ongoing dataset evaluation, ensuring that LLM applications for information retrieval and document analysis continue to meet evolving standards.

연구 동기 및 목표

  • 문서 기반 질의응답 작업에서 정확한 답변 선택 및 수치 추출이 요구되는 대규모 언어 모델(Large Language Models, LLMs)의 성능 평가.
  • 정확한 정보 추출이 필수적인 분야인 의료 메타분석과 같은 실생활 응용 분야에서 LLM의 신뢰성 평가.
  • 질의 형식, 특히 다중 선택 및 수치 추출이 LLM 성능에 미치는 영향 조사.
  • 전문가가 주석 처리한 사실 기반 응답을 제공하는 CogTale 데이터셋을 활용하여 벤치마크 수립.
  • 고위험·정밀도 중심의 문서 분석 작업에 투입될 때 LLM 능력의 격차 규명

제안 방법

  • 연구는 노인 환자 대상 공개된 인지 간섭 연구에서 유래한 구조적이고 전문가가 주석 처리한 정보를 담고 있는 CogTale 데이터셋을 사용한다.
  • 질의는 네 가지 유형으로 분류된다: 단일 선택, 예/아니요, 다중 선택, 수치 추출. 각 유형은 사전 정의된 선택지 중 정확한 답변 선택을 요구한다.
  • LLM(GPT-4 및 GPT-3.5)은 피팅 조정 없이 직접 프ом프팅을 통해 문서 맥락과 질문을 제공하는 제로샷 설정에서 평가된다.
  • 성능는 질문 유형 간 정확한 일치 정확도를 측정하여 모델 및 형식 간 비교 분석한다.
  • 검색기에서 관련 맥락을 제공한다고 가정하여, LLM의 답변 생성 성능만을 분리 평가한다.
  • 연구는 CogTale 플랫폼에서 선별한 13개의 연구에 대해 수행되어, 문서 품질과 영역 특이성의 일관성을 확보한다.
Figure 1: Question-Answering (QA) Framework using LLM for a document-based QA task.
Figure 1: Question-Answering (QA) Framework using LLM for a document-based QA task.

실험 결과

연구 질문

  • RQ1LLM은 사전 정의된 선택지 중 정확한 답변 선택이 요구되는 문서 기반 QA 작업에서 얼마나 잘 수행되는가?
  • RQ2GPT-4와 GPT-3.5는 단일 선택, 예/아니요, 다중 선택, 수치 추출 질문에서 성능에 어떤 차이를 보이는가?
  • RQ3다중 선택 및 수치 추출과 같은 질의 형식이 제로샷 설정에서 LLM 성능에 얼마나 큰 도전을 가하는가?
  • RQ4임상 및 연구 문서 맥락에서 다양한 유형의 사실 추출 작업에서 LLM 성능는 어떻게 변동하는가?
  • RQ5고도의 프롬프팅 전략이 다중 선택 및 수치 추출과 같은 어려운 질의 형식에서 LLM 성능 향상에 기여하는가?

주요 결과

  • GPT-4는 단일 선택 및 예/아니요 질문에서 높은 정확도를 달성하여 맥락 이해력과 정확한 답변 선택 능력이 뛰어나다는 것을 입증했다.
  • 다중 선택 질문에서는 성능 저하가 심각하게 나타나, 유사한 대안들 중에서 정답을 선택하는 데 어려움을 겪고 있음을 보여주었다.
  • 수치 추출 작업은 가장 낮은 성능를 보였으며, LLM이 문서에서 수치 값을 정확히 식별하거나 추출하지 못하는 경우가 빈번했다.
  • GPT-3.5는 GPT-4에 비해 전반적인 정확도가 낮았으며, 특히 다중 선택 및 수치 추출과 같은 복잡한 형식에서 뚜렷한 성능 격차를 보였다.
  • GPT-4의 성능 향상에도 불구하고, 비용(26.54달러)은 GPT-3.5-turbo(1.77달러)의 15배에 이르며, 실용적 투입 시 비용 대비 효율성에 대한 우려를 제기한다.
  • 본 연구는 정확한 사실 기반 추출이 필수적인 시스템적 리뷰 및 메타분석과 같은 정밀도 중심 응용 분야에서 LLM의 신뢰성에 심각한 격차가 있음을 규명했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.