[Paper Review] GujiBERT and GujiGPT: Construction of Intelligent Information Processing Foundation Language Models for Ancient Texts
The paper presents GujiBERT and GujiGPT, foundation language models tailored for intelligent processing of ancient Chinese texts, covering tasks from segmentation to translation.
In the context of the rapid development of large language models, we have meticulously trained and introduced the GujiBERT and GujiGPT language models, which are foundational models specifically designed for intelligent information processing of ancient texts. These models have been trained on an extensive dataset that encompasses both simplified and traditional Chinese characters, allowing them to effectively handle various natural language processing tasks related to ancient books, including but not limited to automatic sentence segmentation, punctuation, word segmentation, part-of-speech tagging, entity recognition, and automatic translation. Notably, these models have exhibited exceptional performance across a range of validation tasks using publicly available datasets. Our research findings highlight the efficacy of employing self-supervised methods to further train the models using classical text corpora, thus enhancing their capability to tackle downstream tasks. Moreover, it is worth emphasizing that the choice of font, the scale of the corpus, and the initial model selection all exert significant influence over the ultimate experimental outcomes. To cater to the diverse text processing preferences of researchers in digital humanities and linguistics, we have developed three distinct categories comprising a total of nine model variations. We believe that by sharing these foundational language models specialized in the domain of ancient texts, we can facilitate the intelligent processing and scholarly exploration of ancient literary works and, consequently, contribute to the global dissemination of China's rich and esteemed traditional culture in this new era.
Motivation & Objective
- Motivate development of specialized LMs for ancient texts to support digital humanities and linguistic research.
- Create models capable of handling both simplified and traditional Chinese characters.
- Demonstrate performance on tasks such as sentence segmentation, punctuation, word segmentation, POS tagging, entity recognition, and translation.
Proposed method
- Train GujiBERT and GujiGPT on large corpora of ancient/classical Chinese and modern Chinese with both simplified and traditional characters.
- Evaluate across multiple NLP tasks including automatic sentence segmentation, punctuation, word segmentation, POS tagging, entity recognition, and automatic translation.
- Apply self-supervised refinement using classical text corpora to enhance downstream task performance.
- Explore the impact of font choice, corpus scale, and initial model selection on outcomes.
- Provide three categories and nine model variations to accommodate different researcher preferences in digital humanities and linguistics.
Experimental results
Research questions
- RQ1How can foundation language models be specialized for intelligent information processing of ancient texts?
- RQ2What is the impact of font, corpus scale, and initial model choice on performance for ancient Chinese NLP tasks?
- RQ3Can self-supervised fine-tuning on classical corpora improve downstream tasks like segmentation, tagging, NER, and translation?
- RQ4Do multiple model variants cover diverse user needs in digital humanities and linguistics?
Key findings
- GujiBERT and GujiGPT achieve strong performance on a range of ancient text processing tasks.
- Self-supervised training on classical corpora enhances downstream task capabilities.
- Font, corpus scale, and initial model selection significantly influence experimental outcomes.
- Three categories with nine model variations offer flexibility for researchers with different preferences.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.