[论文解读] GujiBERT and GujiGPT: Construction of Intelligent Information Processing Foundation Language Models for Ancient Texts
这篇论文提出 GujiBERT 和 GujiGPT,这些基础语言模型专为古代汉文的智能处理而设计,覆盖从分词到翻译等任务。
In the context of the rapid development of large language models, we have meticulously trained and introduced the GujiBERT and GujiGPT language models, which are foundational models specifically designed for intelligent information processing of ancient texts. These models have been trained on an extensive dataset that encompasses both simplified and traditional Chinese characters, allowing them to effectively handle various natural language processing tasks related to ancient books, including but not limited to automatic sentence segmentation, punctuation, word segmentation, part-of-speech tagging, entity recognition, and automatic translation. Notably, these models have exhibited exceptional performance across a range of validation tasks using publicly available datasets. Our research findings highlight the efficacy of employing self-supervised methods to further train the models using classical text corpora, thus enhancing their capability to tackle downstream tasks. Moreover, it is worth emphasizing that the choice of font, the scale of the corpus, and the initial model selection all exert significant influence over the ultimate experimental outcomes. To cater to the diverse text processing preferences of researchers in digital humanities and linguistics, we have developed three distinct categories comprising a total of nine model variations. We believe that by sharing these foundational language models specialized in the domain of ancient texts, we can facilitate the intelligent processing and scholarly exploration of ancient literary works and, consequently, contribute to the global dissemination of China's rich and esteemed traditional culture in this new era.
研究动机与目标
- 激励为古文本开发专门的语言模型,以支持数字人文和语言学研究。
- 构建能够处理简体和繁体汉字的模型。
- 展示在句子分割、标点、分词、词性标注、实体识别和翻译等任务上的性能。
提出的方法
- 在包含古代/经典汉语和现代汉语的庞大语料库上训练 GujiBERT 和 GujiGPT,涵盖简体和繁体字符。
- 在多种NLP任务上进行评估,包括自动句子分割、标点、分词、词性标注、实体识别和自动翻译。
- 使用经典文本语料进行自监督微调,以提升下游任务表现。
- 探索字体选择、语料规模和初始模型选择对结果的影响。
- 提供三大类和九种模型变体,以满足数字人文与语言学研究者的不同偏好。
实验结果
研究问题
- RQ1基础语言模型如何针对古文本的智能信息处理进行专业化?
- RQ2字体、语料规模和初始模型选择对古代汉语NLP任务的性能有何影响?
- RQ3在经典语料上进行自监督微调是否能改进下游任务,如分割、标注、NER 和翻译?
- RQ4多种模型变体是否能覆盖数字人文与语言学中多样的用户需求?
主要发现
- GujiBERT 和 GujiGPT 在一系列古文本处理任务上取得了出色的表现。
- 在经典语料上进行自监督训练提升了下游任务能力。
- 字体、语料规模和初始模型选择显著影响实验结果。
- 三大类九种模型变体为具有不同偏好的研究者提供了灵活性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。