Skip to main content
QUICK REVIEW

[论文解读] An Empirical Study on Large Language Models in Accuracy and Robustness under Chinese Industrial Scenarios

Zongjie Li, Wenying Qiu|arXiv (Cornell University)|Jan 27, 2024
Natural Language Processing Techniques被引用 4
一句话总结

本文通过一项全面的实证研究,评估了10个大型语言模型(9个中文本地模型和4个全球模型)在8个中文工业领域中的准确性和鲁棒性。研究基于人工整理的1,200个领域特定问题基准,以及涵盖8项能力的13,631个变体问题的元测试框架,发现所有LLM的准确率均低于0.6,其中全球模型在推理和开放式任务中表现优于本地模型,而本地模型在中文工业术语理解方面表现更优。

ABSTRACT

Recent years have witnessed the rapid development of large language models (LLMs) in various domains. To better serve the large number of Chinese users, many commercial vendors in China have adopted localization strategies, training and providing local LLMs specifically customized for Chinese users. Furthermore, looking ahead, one of the key future applications of LLMs will be practical deployment in industrial production by enterprises and users in those sectors. However, the accuracy and robustness of LLMs in industrial scenarios have not been well studied. In this paper, we present a comprehensive empirical study on the accuracy and robustness of LLMs in the context of the Chinese industrial production area. We manually collected 1,200 domain-specific problems from 8 different industrial sectors to evaluate LLM accuracy. Furthermore, we designed a metamorphic testing framework containing four industrial-specific stability categories with eight abilities, totaling 13,631 questions with variants to evaluate LLM robustness. In total, we evaluated 9 different LLMs developed by Chinese vendors, as well as four different LLMs developed by global vendors. Our major findings include: (1) Current LLMs exhibit low accuracy in Chinese industrial contexts, with all LLMs scoring less than 0.6. (2) The robustness scores vary across industrial sectors, and local LLMs overall perform worse than global ones. (3) LLM robustness differs significantly across abilities. Global LLMs are more robust under logical-related variants, while advanced local LLMs perform better on problems related to understanding Chinese industrial terminology. Our study results provide valuable guidance for understanding and promoting the industrial domain capabilities of LLMs from both development and industrial enterprise perspectives. The results further motivate possible research directions and tooling support.

研究动机与目标

  • 评估大型语言模型(LLMs)在真实中文工业生产环境中的准确性和鲁棒性,因为以往的基准测试多集中于对话或非工业任务。
  • 解决中文工业LLM缺乏领域特定评估框架的问题,尤其是在本地化模型日益针对中文用户定制的背景下。
  • 比较中文本地化LLM与全球开发LLM在不同工业领域中,于准确性、鲁棒性和领域特定能力方面的表现差异。
  • 为工业企业和模型开发者提供关于模型选择、配置以及未来工业AI部署工具需求的可操作洞察。
  • 建立一个可复现、可扩展的评估框架,用于评估工业场景中的LLM,该框架可推广至其他语言和领域。

提出的方法

  • 从8个中文工业领域中人工整理出1,200个行业特定问题的基准,涵盖多样的技术和运营挑战。
  • 设计了一种元测试框架,包含四个工业特定的稳定性类别和八项能力,生成13,631个变体问题,以评估在语义扰动下的鲁棒性。
  • 在标准化条件下评估10个LLM(9个由中国厂商开发,4个为全球模型),温度设置为0以获得确定性输出。
  • 应用领域特定的元关系测试输入变化下的一致性,如同义词替换、改写和逻辑转换。
  • 结合人工标注的正确答案与自动化评分,测量各项任务中的准确性和鲁棒性。
  • 在各领域开展大规模实验,比较模型表现,重点关注推理、术语理解以及在扰动下的稳定性差异。

实验结果

研究问题

  • RQ1中文本地化模型与全球开发LLM在中文工业场景的真实工业问题上的准确率表现如何比较?
  • RQ2LLM在与工业生产任务相关的语义变化下,其鲁棒性在多大程度上得以保持?
  • RQ3模型表现如何随不同工业领域而变化?哪些能力对扰动最为敏感?
  • RQ4本地LLM与全球LLM在理解中文工业术语与逻辑推理及开放式生成方面,各自的优势是什么?
  • RQ5元测试框架能否有效评估工业场景中LLM的鲁棒性?其在不同领域和语言间的通用性如何?

主要发现

  • 所有评估的LLM在中文工业基准上的准确率均低于0.6,表明其在工业部署中仍有巨大提升空间。
  • 全球LLM在推理和开放式生成任务中表现优于本地LLM,而本地LLM在理解中文工业术语方面表现出更强能力。
  • 鲁棒性在不同能力间差异显著:全球模型在逻辑和句法变体下表现出更强的抗扰能力,而先进本地模型在术语相关扰动中表现更优。
  • 总体而言,本地LLM在所有工业领域中的鲁棒性得分均低于全球LLM,凸显其在输入变化下的稳定性差距。
  • 元测试框架成功识别出不同模型和能力之间的性能差异与敏感性,证明其在工业LLM评估中的实用性。
  • 模型表现对配置设置敏感,固定温度(0)可能低估真实场景下的有效性,提示未来评估需依赖厂商指导的校准。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。