Skip to main content
QUICK REVIEW

[논문 리뷰] Multilingual Tourist Assistance using ChatGPT: Comparing Capabilities in Hindi, Telugu, and Kannada

Sanjana Kolar, Rohit Kumar|arXiv (Cornell University)|2023. 07. 28.
Artificial Intelligence in Healthcare and Education인용 수 8
한 줄 요약

이 논문은 50문항 BLEU 기반 평가와 인간 평가 연구를 통해 ChatGPT의 영어-인도 어언어(힌디어, 칸나다, 텔루구) 번역을 평가하고, 힌디어가 전반적으로 우수하다고 결론내립니다.

ABSTRACT

This research investigates the effectiveness of ChatGPT, an AI language model by OpenAI, in translating English into Hindi, Telugu, and Kannada languages, aimed at assisting tourists in India's linguistically diverse environment. To measure the translation quality, a test set of 50 questions from diverse fields such as general knowledge, food, and travel was used. These were assessed by five volunteers for accuracy and fluency, and the scores were subsequently converted into a BLEU score. The BLEU score evaluates the closeness of a machine-generated translation to a human translation, with a higher score indicating better translation quality. The Hindi translations outperformed others, showcasing superior accuracy and fluency, whereas Telugu translations lagged behind. Human evaluators rated both the accuracy and fluency of translations, offering a comprehensive perspective on the language model's performance.

연구 동기 및 목표

  • 인도에서의 관광 정보에 대해 영어를 힌디어, 칸나다, 텔루구로 번역하는 ChatGPT의 번역 품질을 평가한다.
  • 주관적(정확도 및 유창도 평가) 및 객관적(BLEU 점수) 측정을 사용하여 번역 품질을 정량화한다.
  • 관광 도메인 번역에 대한 언어별 강점과 약점을 식별하여 개선 방향을 제시한다.

제안 방법

  • 시스템 역할 + 사용자 역할의 이중 프롬트를 통해 영어 텍스트를 대상 인도 언어로 번역하기 위해 gpt-3.5-turbo를 사용한다.
  • 정확도와 유창도를 5명의 원어민 자원봉사자에게 평가한다(척도 1-5).
  • 50문항에 대해 기계 번역과 참조 인간 번역을 비교하여 BLEU 점수(0-100)를 계산한다.
  • 초기 60문항을 3개의 주제(일반, 음식, 여행)로 분류하고 50개 관련 항목을 선택한다.

실험 결과

연구 질문

  • RQ1챗GPT가 독일/국제 방문객을 대상으로 영어 관광 질의를 힌디, 칸나다, 텔루구로 정확하게 번역할 수 있는가?
  • RQ2세 언어 간 주관적(정확도/유창도) 평가와 객관적(BLEU) 평가의 차이가 있는가?
  • RQ3관광 도메인 번역을 향상시키기 위한 언어별 구체적 개선은 무엇인가?

주요 결과

  • 히디어 번역은 전반적으로 정확도와 유창도가 가장 높다(예: 일반 영역: 정확도 4.8, 유창도 4.6).
  • 텔루구 번역은 가장 낮은 성능을 보인다(BLEU: 13.12; 일반 정확도: 2.6; 유창도: 2.1).
  • 칸나다 번역은 중간 수준이다(BLEU: 46.78; 일반 정확도: 3.7; 유창도: 3.5).
  • 전반적으로 힌디어 일반/주제 번역이 칸나다와 텔루구보다 정확도와 유창도 면에서 더 높게 나타난다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.