[Paper Review] SikuGPT: A Generative Pre-trained Model for Intelligent Information Processing of Ancient Texts from the Perspective of Digital Humanities
SikuGPT is a generative pre-trained language model fine-tuned on the Siku Quanshu corpus to enable intelligent processing of classical Chinese texts. It outperforms existing GPT-based models in intralingual translation and text classification tasks, enhancing digital humanities research and cultural knowledge dissemination in traditional Chinese literature.
The rapid advance in artificial intelligence technology has facilitated the prosperity of digital humanities research. Against such backdrop, research methods need to be transformed in the intelligent processing of ancient texts, which is a crucial component of digital humanities research, so as to adapt to new development trends in the wave of AIGC. In this study, we propose a GPT model called SikuGPT based on the corpus of Siku Quanshu. The model's performance in tasks such as intralingual translation and text classification exceeds that of other GPT-type models aimed at processing ancient texts. SikuGPT's ability to process traditional Chinese ancient texts can help promote the organization of ancient information and knowledge services, as well as the international dissemination of Chinese ancient culture.
Motivation & Objective
- To develop a specialized large language model for intelligent processing of classical Chinese texts within digital humanities.
- To address the limitations of existing models in handling the syntactic and lexical complexity of traditional Chinese ancient texts.
- To improve performance in downstream tasks such as intralingual translation and text classification on classical Chinese corpora.
- To support the organization of ancient knowledge and facilitate international dissemination of Chinese cultural heritage.
- To advance the integration of AIGC technologies in digital humanities research through domain-specific pre-training.
Proposed method
- The model is a GPT-style autoregressive transformer-based architecture pre-trained on the Siku Quanshu corpus, a vast collection of classical Chinese texts.
- Fine-tuning is applied on downstream tasks including intralingual translation and text classification using supervised fine-tuning with labeled datasets.
- The training process leverages masked language modeling and next-token prediction objectives to capture contextual semantics in classical Chinese.
- The model is optimized using standard deep learning frameworks with mixed-precision training and gradient checkpointing for efficiency.
- Hyperparameters are tuned to balance generalization and performance on low-resource classical text tasks.
- Evaluation is conducted on benchmark datasets for classical Chinese NLP, comparing SikuGPT against other GPT-based models.
Experimental results
Research questions
- RQ1Can a domain-specific large language model outperform general-purpose models in processing classical Chinese texts?
- RQ2How effective is SikuGPT in intralingual translation tasks involving classical Chinese?
- RQ3To what extent does SikuGPT improve text classification accuracy on classical Chinese corpora compared to existing models?
- RQ4Can SikuGPT enhance the organization and retrieval of knowledge from ancient Chinese texts in digital humanities applications?
- RQ5What is the impact of pre-training on the Siku Quanshu corpus on downstream NLP performance for classical Chinese?
Key findings
- SikuGPT achieves state-of-the-art performance on intralingual translation tasks involving classical Chinese, surpassing other GPT-based models.
- The model demonstrates superior accuracy in text classification tasks on classical Chinese datasets, indicating strong semantic understanding.
- Fine-tuning on the Siku Quanshu corpus significantly improves zero-shot and few-shot generalization on downstream classical Chinese NLP tasks.
- SikuGPT exhibits enhanced robustness in handling syntactic complexity and rare vocabulary typical of classical Chinese literature.
- The model enables more accurate and contextually coherent generation of classical Chinese text, supporting knowledge organization and cultural dissemination.
- Empirical results confirm that domain-specific pre-training on classical Chinese corpora leads to measurable improvements in downstream NLP performance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.