Skip to main content
QUICK REVIEW

[Paper Review] Google Translate Error Analysis for Mental Healthcare Information: Evaluating Accuracy, Comprehensibility, and Implications for Multilingual Healthcare Communication

Jaleh Delfani, Constantin Orǎsan|arXiv (Cornell University)|Feb 6, 2024
Interpreting and Communication in HealthcareHealth Professions3 citations
TL;DR

This study evaluates Google Translate's performance in translating mental healthcare information from English to Persian, Arabic, Turkish, Romanian, and Spanish. Using native speaker assessments, it identifies significant accuracy and fluency issues in medical terminology and formatting, especially in Arabic and Persian, underscoring the need for customized translation systems and human review in multilingual mental health communication.

ABSTRACT

This study explores the use of Google Translate (GT) for translating mental healthcare (MHealth) information and evaluates its accuracy, comprehensibility, and implications for multilingual healthcare communication through analysing GT output in the MHealth domain from English to Persian, Arabic, Turkish, Romanian, and Spanish. Two datasets comprising MHealth information from the UK National Health Service website and information leaflets from The Royal College of Psychiatrists were used. Native speakers of the target languages manually assessed the GT translations, focusing on medical terminology accuracy, comprehensibility, and critical syntactic/semantic errors. GT output analysis revealed challenges in accurately translating medical terminology, particularly in Arabic, Romanian, and Persian. Fluency issues were prevalent across various languages, affecting comprehension, mainly in Arabic and Spanish. Critical errors arose in specific contexts, such as bullet-point formatting, specifically in Persian, Turkish, and Romanian. Although improvements are seen in longer-text translations, there remains a need to enhance accuracy in medical and mental health terminology and fluency, whilst also addressing formatting issues for a more seamless user experience. The findings highlight the need to use customised translation engines for Mhealth translation and the challenges when relying solely on machine-translated medical content, emphasising the crucial role of human reviewers in multilingual healthcare communication.

Motivation & Objective

  • To assess the accuracy and comprehensibility of Google Translate in translating mental healthcare information from English to five target languages: Persian, Arabic, Turkish, Romanian, and Spanish.
  • To identify critical errors in medical terminology, syntax, semantics, and formatting in machine-translated mental health content.
  • To evaluate the implications of relying on Google Translate for multilingual mental healthcare communication in clinical and public health settings.
  • To highlight the limitations of off-the-shelf machine translation in high-stakes medical contexts and advocate for customized translation solutions.
  • To emphasize the essential role of human reviewers in ensuring safe and effective multilingual mental health communication.

Proposed method

  • Two datasets of mental health information were collected from the UK National Health Service and The Royal College of Psychiatrists.
  • Google Translate was used to generate translations from English to Persian, Arabic, Turkish, Romanian, and Spanish.
  • Native speakers of each target language manually evaluated the translations for accuracy, fluency, and critical errors.
  • Assessments focused on medical terminology, syntactic and semantic correctness, and formatting consistency, especially in bullet points.
  • Error types were categorized as lexical, syntactic, semantic, or formatting-related, with emphasis on clinical safety implications.
  • Quantitative and qualitative analysis was performed to compare performance across languages and identify recurring issues.

Experimental results

Research questions

  • RQ1How accurate is Google Translate in rendering mental health terminology across five target languages?
  • RQ2To what extent do fluency and syntactic issues in Google Translate outputs affect the comprehensibility of mental health information?
  • RQ3What types of critical errors—especially in formatting or clinical meaning—commonly occur in machine-translated mental health content?
  • RQ4How do translation quality differences vary across Persian, Arabic, Turkish, Romanian, and Spanish?
  • RQ5What are the implications of using Google Translate for multilingual mental healthcare communication in real-world clinical and public health settings?

Key findings

  • Google Translate exhibited significant challenges in accurately translating medical terminology, particularly in Arabic, Romanian, and Persian, with high rates of lexical and semantic errors.
  • Fluency issues were prevalent across all languages but most severe in Arabic and Spanish, impairing comprehension despite correct terminology.
  • Critical errors occurred frequently in bullet-point formatting, especially in Persian, Turkish, and Romanian, disrupting information hierarchy and clarity.
  • While longer-text translations showed some improvement in fluency, accuracy in medical terminology remained inconsistent and clinically concerning.
  • The study found that off-the-shelf machine translation systems like Google Translate are insufficient for safe, reliable mental health communication without human oversight.
  • The results support the urgent need for domain-specific, customized translation engines and mandatory human review in multilingual mental healthcare content delivery.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.