Skip to main content
QUICK REVIEW

[论文解读] Prompting Multilingual Large Language Models to Generate Code-Mixed Texts: The Case of South East Asian Languages

Zheng-Xin Yong, Ruochen Zhang|arXiv (Cornell University)|Mar 23, 2023
Natural Language Processing Techniques被引用 6
一句话总结

本文研究了多语言大型语言模型(LLMs)在零样本提示下生成七种东南亚语言(印度尼西亚语、马来语、汉语、他加禄语、越南语、泰米尔语和新加坡英语)自然流畅的代码混用文本的能力。研究发现,尽管像ChatGPT这样的模型在特定语言对(如英语-新加坡英语)中表现尚可,但大多数模型在跨不同语系生成时,难以生成语法正确或语义有意义的代码混用文本,且常引入非预期语言或错误的书写系统,凸显了在自然语言处理研究中使用前进行广泛人工验证的必要性。

ABSTRACT

While code-mixing is a common linguistic practice in many parts of the world, collecting high-quality and low-cost code-mixed data remains a challenge for natural language processing (NLP) research. The recent proliferation of Large Language Models (LLMs) compels one to ask: how capable are these systems in generating code-mixed data? In this paper, we explore prompting multilingual LLMs in a zero-shot manner to generate code-mixed data for seven languages in South East Asia (SEA), namely Indonesian, Malay, Chinese, Tagalog, Vietnamese, Tamil, and Singlish. We find that publicly available multilingual instruction-tuned models such as BLOOMZ and Flan-T5-XXL are incapable of producing texts with phrases or clauses from different languages. ChatGPT exhibits inconsistent capabilities in generating code-mixed texts, wherein its performance varies depending on the prompt template and language pairing. For instance, ChatGPT generates fluent and natural Singlish texts (an English-based creole spoken in Singapore), but for English-Tamil language pair, the system mostly produces grammatically incorrect or semantically meaningless utterances. Furthermore, it may erroneously introduce languages not specified in the prompt. Based on our investigation, existing multilingual LLMs exhibit a wide range of proficiency in code-mixed data generation for SEA languages. As such, we advise against using LLMs in this context without extensive human checks.

研究动机与目标

  • 评估零样本提示在资源匮乏的东南亚语言中生成高质量、自然流畅的代码混用文本的可行性。
  • 识别当前多语言LLM在跨多样化语言对生成语言准确且语义连贯的代码混用语句方面的局限性。
  • 评估LLM在生成用于自然语言处理研究的合成代码混用数据时的可靠性,尤其考虑到该领域高质量标注数据集的稀缺性。
  • 为ChatGPT、BLOOMZ和Flan-T5-XXL等模型在不同语系和代码混用层级下的代码混用生成表现提供实证证据。

提出的方法

  • 对五种多语言LLM(ChatGPT、InstructGPT、BLOOMZ和Flan-T5-XXL)进行零样本提示,使用多样化的提示模板,生成与六种东南亚语言配对的英语代码混用文本。
  • 设计模拟双语使用者或模仿特定语体风格的提示模板,以激发代码混用响应。
  • 通过母语者对模型输出进行收集与标注,从多个维度评估自然度与代码混用程度:语法正确性、语义连贯性及书写系统准确性。
  • 将代码混用程度划分为四类:无代码混用、借词、主题相关名词及完整语言元素。
  • 对错误进行定性与定量分析,包括语法错误、语义混淆、书写系统误用及非预期语言的引入。
  • 比较不同语言对与代码混用类型下的模型输出,识别失败模式与性能差异的规律。
Figure 1: Depiction of SEA regions, which consist of a total of 11 countries. We prompt LLMs to generate code-mixed data of languages used in six South East Asian countries (colored in dark blue): Brunei, Indonesia, Malaysia, Philippines, Singapore, and Vietnam.
Figure 1: Depiction of SEA regions, which consist of a total of 11 countries. We prompt LLMs to generate code-mixed data of languages used in six South East Asian countries (colored in dark blue): Brunei, Indonesia, Malaysia, Philippines, Singapore, and Vietnam.

实验结果

研究问题

  • RQ1多语言LLM在零样本设置下,对资源匮乏的东南亚语言,能在多大程度上生成自然流畅的代码混用文本?
  • RQ2不同LLM(如ChatGPT、BLOOMZ、Flan-T5-XXL)在跨多样化语言对生成语法正确且语义明确的代码混用语句方面的能力有何差异?
  • RQ3LLM生成的代码混用文本中最常见的错误类型是什么,例如语法错误、语义不连贯或书写系统不匹配?
  • RQ4为何某些语言对(如英语-新加坡英语)能生成更流畅的输出,而另一些(如英语-泰米尔语)则不能?哪些语言或结构因素导致这种差异?
  • RQ5LLM在输出中引入非指定语言或未能遵守指定语言对的情况有多严重?

主要发现

  • ChatGPT仅在特定语言对(如英语-新加坡英语)中能生成流畅自然的代码混用文本,但在其他语言对(如英语-泰米尔语)中表现较差,输出常出现语法错误或语义无意义。
  • BLOOMZ和Flan-T5-XXL基本无法生成任何形式的代码混用文本,即使使用强烈提示,也仅产生单语或非习惯表达。
  • 许多模型输出存在严重的自然度问题,包括不自然的表达、错误的动词形式、误用所有格标记,以及语义矛盾,例如声称已抵达却仍被困在交通中。
  • 大量输出存在书写系统不一致问题,如将拉丁转写与原生书写系统混合使用(如泰米尔字母与拉丁字母混用),或在中英代码混用中使用汉语拼音而非汉字。
  • LLM频繁引入提示中未指定的语言,如福建话或其他方言,导致输出中出现非预期的语言污染。
  • LLM在不同语系间的性能差异显著,涉及德拉维达语系(如泰米尔语)和南亚语系(如越南语)的语言对中错误率更高,而与南岛语系或汉藏语系语言对相比则表现更优。
(a) Template: Assume to be bilingual speaker
(a) Template: Assume to be bilingual speaker

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。