Skip to main content
QUICK REVIEW

[논문 리뷰] GujiBERT and GujiGPT: Construction of Intelligent Information Processing Foundation Language Models for Ancient Texts

Dongbo Wang, Chang Liu|arXiv (Cornell University)|2023. 07. 11.
Computational and Text Analysis Methods인용 수 11
한 줄 요약

본 논문은 GujiBERT와 GujiGPT를 제시하며, 고대 중국어 텍스트의 지능적 처리를 위해 맞춤화된 기본 언어 모델로, 구분에서 번역에 이르는 태스크를 다룬다.

ABSTRACT

In the context of the rapid development of large language models, we have meticulously trained and introduced the GujiBERT and GujiGPT language models, which are foundational models specifically designed for intelligent information processing of ancient texts. These models have been trained on an extensive dataset that encompasses both simplified and traditional Chinese characters, allowing them to effectively handle various natural language processing tasks related to ancient books, including but not limited to automatic sentence segmentation, punctuation, word segmentation, part-of-speech tagging, entity recognition, and automatic translation. Notably, these models have exhibited exceptional performance across a range of validation tasks using publicly available datasets. Our research findings highlight the efficacy of employing self-supervised methods to further train the models using classical text corpora, thus enhancing their capability to tackle downstream tasks. Moreover, it is worth emphasizing that the choice of font, the scale of the corpus, and the initial model selection all exert significant influence over the ultimate experimental outcomes. To cater to the diverse text processing preferences of researchers in digital humanities and linguistics, we have developed three distinct categories comprising a total of nine model variations. We believe that by sharing these foundational language models specialized in the domain of ancient texts, we can facilitate the intelligent processing and scholarly exploration of ancient literary works and, consequently, contribute to the global dissemination of China's rich and esteemed traditional culture in this new era.

연구 동기 및 목표

  • 고전 텍스트를 위한 전문화된 LM의 개발을 촉진해 디지털 인문학 및 언어학 연구를 지원한다.
  • 간체자와 번체자를 모두 처리할 수 있는 모델을 만든다.
  • 문장 구분, 구두점 표시, 어절 분할, 품사 태깅, 엔터티 인식, 번역 등의 태스크에서 성능을 보여준다.

제안 방법

  • 간체자와 번체자를 모두 포함한 고대/전통 중국어와 현대 중국어의 대규모 말뭉치로 GujiBERT와 GujiGPT를 학습한다.
  • 자동 문장 구분, 구두점, 어절 분할, 품사 태깅, 엔터티 인식, 자동 번역 등 여러 NLP 태스크를 평가한다.
  • 고전 문자 말뭉치를 이용한 자기지도 정제를 적용해 다운스트림 태스크 성능을 향상시킨다.
  • 폰트 선택, 말뭉치 규모, 초기 모델 선택이 결과에 미치는 영향을 탐구한다.
  • 디지털 인문학 및 언어학에서 다양한 연구자 선호를 충족시키기 위해 세 가지 범주와 아홉 가지 모델 변형을 제공한다.

실험 결과

연구 질문

  • RQ1기초 언어 모델을 어떻게 고대 텍스트의 지능적 정보 처리에 특화시킬 수 있는가?
  • RQ2고대 중국어 NLP 태스크의 성능에 있어 폰트, 말뭉치 규모, 초기 모델 선택의 영향은 무엇인가?
  • RQ3고전 말뭉치에 대한 자기지도 미세조정이 구분, 태깅, NER, 번역과 같은 다운스트림 작업을 향상시킬 수 있는가?
  • RQ4다양한 모델 변형이 디지털 인문학과 언어학의 다양한 사용자 요구를 충족하는가?

주요 결과

  • GujiBERT와 GujiGPT는 다양한 고대 텍스트 처리 태스크에서 강력한 성능을 달성한다.
  • 고전 말뭉치를 활용한 자기지도 학습은 다운스트림 태스크 역량을 강화한다.
  • 폰트, 말뭉치 규모, 초기 모델 선택은 실험 결과에 상당한 영향을 미친다.
  • 세 가지 범주와 아홉 가지 모델 변형은 서로 다른 선호를 가진 연구자들에게 유연성을 제공한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.