[논문 리뷰] Is ChatGPT a Good NLG Evaluator? A Preliminary Study
본 논문은 일반 NLG 평가 지표로서의 ChatGPT를 예비적으로 평가하고, 프롬프트와 데이터 세트 편향의 영향으로 인간 판단에 대한 다수의 메타 평가 데이터 세트에서 높은 상관관계를 보이며 성능이 나타난다는 것을 발견했다.
Recently, the emergence of ChatGPT has attracted wide attention from the computational linguistics community. Many prior studies have shown that ChatGPT achieves remarkable performance on various NLP tasks in terms of automatic evaluation metrics. However, the ability of ChatGPT to serve as an evaluation metric is still underexplored. Considering assessing the quality of natural language generation (NLG) models is an arduous task and NLG metrics notoriously show their poor correlation with human judgments, we wonder whether ChatGPT is a good NLG evaluation metric. In this report, we provide a preliminary meta-evaluation on ChatGPT to show its reliability as an NLG metric. In detail, we regard ChatGPT as a human evaluator and give task-specific (e.g., summarization) and aspect-specific (e.g., relevance) instruction to prompt ChatGPT to evaluate the generated results of NLG models. We conduct experiments on five NLG meta-evaluation datasets (including summarization, story generation and data-to-text tasks). Experimental results show that compared with previous automatic metrics, ChatGPT achieves state-of-the-art or competitive correlation with human judgments in most cases. In addition, we find that the effectiveness of the ChatGPT evaluator might be influenced by the creation method of the meta-evaluation datasets. For the meta-evaluation datasets which are created greatly depending on the reference and thus are biased, the ChatGPT evaluator might lose its effectiveness. We hope our preliminary study could prompt the emergence of a general-purposed reliable NLG metric.
연구 동기 및 목표
- ChatGPT가 일반 NLG 평가 지표로서 기능할 수 있는지 동기 부여 및 평가.
- 전통적인 자동 평가 지표의 한계와 참조-무 의 평가자 및 참조 기반 평가자로서의 ChatGPT의 가능성 제시.
- 작업별 및 측면별 프롬프트가 요약, 이야기 생성, 데이터-텍스트 변환 등 NLG 출력에 대한 ChatGPT의 판단에 미치는 영향 조사.
- 데이터 세트 구성 편향이 ChatGPT를 평가 지표로서의 효과에 미치는 영향 검토.
제안 방법
- ChatGPT를 인간 평가자로 간주하고 작업별 및 측면별 프롬프트를 적용하여 점수나 등급을 산출.
- ChatGPT 기반 평가를 표준 자동 지표(ROUGE, BERTScore, MoverScore, PRISM, BARTScore 등)와 비교.
- 참조-무 프롬프트(DA 및 별점)와 참조 기반 프롬프트(황금 참조 포함)를 사용하여 채점 가이드를 제시.
- 요약, 이야기 생성, 데이터-텍스트 변환 작업에 걸친 다섯 가지 메타 평가 데이터 세트에서 평가.
- 샘플 수준 및 데이터 세트 수준의 지표(예: Spearman, Pearson, Kendall)를 사용하여 인간 판단과의 상관관계 분석.
- 프롬프트 설계 및 데이터 세트 구성 편향이 ChatGPT 평가자 성능에 미치는 영향 평가.
실험 결과
연구 질문
- RQ1ChatGPT가 여러 작업에서 인간 판단과의 상관관계를 보이는 일반 NLG 평가자로서 가능할까?
- RQ2작업별 및 측면별 프롬프트가 ChatGPT의 평가 신뢰도에 어떤 영향을 미치는가?
- RQ3메타 평가 데이터 세트의 편향이 ChatGPT를 NLG 지표로서의 효과성에 영향을 주는가?
- RQ4요약, 이야기 생성, 데이터-텍스트 변환 작업에서 다른 확립된 자동 지표와 비교하여 ChatGPT의 성능은 어떤가?
주요 결과
- ChatGPT는 이야기 생성 및 요약 맥락에서 특히 다수의 메타 평가 데이터 세트에서 인간 판단과의 높은 상관관계를 보인다.
- ChatGPT는 여러 작업에서 인간 판단에 대한 상관관계 측면에서 전통적인 자동 지표를 능가하는 경우가 많아 일반 NLG 지표로서의 잠재력을 시사한다.
- 평가자의 효과는 프롬프트 설계에 민감하며, 서로 다른 작업과 측면은 맞춤형 프롬프트를 필요로 한다.
- 참조 유사성에 편향된 데이터 세트는 ChatGPT의 효과를 감소시킬 수 있는데, 어휘적 편향이 참조 기반 신호를 선호하게 만들 수 있다.
- ChatGPT는 요약 및 이야기 생성을 넘어 데이터-텍스트 평가에서도 경쟁력 있는 성능을 보이며 광범위한 적용 가능성을 시사한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.