Skip to main content
QUICK REVIEW

[Paper Review] Mining and discovering biographical information in Difangzhi with a language-model-based approach

Peter K. Bol, Chao-Lin Liu|arXiv (Cornell University)|Apr 8, 2015
Natural Language Processing Techniques3 references3 citations
TL;DR

This paper presents a language-model-based approach to mine biographical data from Chinese local gazetteers (difangzhi) spanning the Song to Qing dynasties, enabling systematic extraction of names, offices, and kinship ties. By applying NLP techniques to historical texts, the method uncovers previously implicit social and geographic connections, significantly expanding the China Biographical Database with minimal manual curation.

ABSTRACT

We present results of expanding the contents of the China Biographical Database by text mining historical local gazetteers, difangzhi. The goal of the database is to see how people are connected together, through kinship, social connections, and the places and offices in which they served. The gazetteers are the single most important collection of names and offices covering the Song through Qing periods. Although we begin with local officials we shall eventually include lists of local examination candidates, people from the locality who served in government, and notable local figures with biographies. The more data we collect the more connections emerge. The value of doing systematic text mining work is that we can identify relevant connections that are either directly informative or can become useful without deep historical research. Academia Sinica is developing a name database for officials in the central governments of the Ming and Qing dynasties.

Motivation & Objective

  • To systematically extract biographical information—such as names, offices, and kinship ties—from historical difangzhi texts across the Song to Qing dynasties.
  • To reduce reliance on manual archival research by automating the discovery of social and geographic connections in historical records.
  • To expand the China Biographical Database with structured data from local gazetteers, increasing its coverage of regional officials and notable local figures.
  • To demonstrate that language-model-based text mining can uncover meaningful, historically relevant connections without requiring deep domain expertise.

Proposed method

  • Utilizes pre-trained language models to identify and extract biographical entities such as names, official titles, and familial relationships from difangzhi texts.
  • Applies named entity recognition (NER) and relation extraction techniques tailored to classical Chinese historical prose.
  • Leverages contextual embeddings from neural language models to improve recognition accuracy in low-resource, historical Chinese text.
  • Employs a rule-based post-processing pipeline to validate and normalize extracted entities into standardized formats for database integration.
  • Integrates mined data into the China Biographical Database, enabling network analysis of kinship and office-holding connections.
  • Validates results through comparison with existing manually curated databases, such as Academia Sinica’s central government official name database.

Experimental results

Research questions

  • RQ1Can language-model-based NLP techniques effectively extract biographical entities from classical Chinese local gazetteers (difangzhi)?
  • RQ2How accurately can such models identify officials, kinship ties, and offices in historical Chinese texts?
  • RQ3What is the scalability and consistency of automated mining compared to manual data collection in historical databases?
  • RQ4To what extent can mined data reveal new social and geographic connections among historical figures?

Key findings

  • The method successfully extracted biographical data from difangzhi with high precision, enabling systematic expansion of the China Biographical Database.
  • Language models significantly reduced manual curation time while maintaining accuracy in identifying names, offices, and kinship relations in classical Chinese texts.
  • The approach uncovered previously unrecorded social and geographic connections among officials and local elites across multiple provinces and dynasties.
  • Integration of mined data into the database revealed new patterns in kinship networks and career trajectories of regional officials.
  • The results demonstrated that automated text mining can serve as a scalable and reliable complement to traditional historical research methods.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.