[Paper Review] Network analysis of named entity interactions in written texts
This paper proposes a network model that links named entities (e.g., characters, locations, organizations) co-occurring in the same textual context to reveal topological structures in written texts. By analyzing books, the model reveals short path lengths, high clustering, and modular organization, outperforming traditional word adjacency networks in identifying unknown references.
The use of methods borrowed from statistics and physics has allowed for the discovery of unprecedent patterns of human behavior and cognition by establishing links between models features and language structure. While current models have been useful to identify patterns via analysis of syntactical and semantical networks, only a few works have probed the relevance of investigating the structure arising from the relationship between relevant entities such as characters, locations and organizations. In this study, we introduce a model that links entities appearing in the same context in order to capture the complexity of entities organization through a networked representation. Computational simulations in books revealed that the proposed model displays interesting topological features, such as short typical shortest path length, high values of clustering coefficient and modular organization. The effectiveness of the our model was verified in a practical pattern recognition task in real networks. When compared with the traditional word adjacency networks, our model displayed optimized results in identifying unknown references in texts. Because the proposed model plays a complementary role in characterizing unstructured documents via topological analysis of named entities, we believe that it could be useful to improve the characterization written texts when combined with other traditional approaches based on statistical and deeper paradigms.
Motivation & Objective
- To investigate the structural organization of named entities in written texts using network analysis.
- To address the limitation of existing models that focus on syntactic or semantic networks rather than entity interactions.
- To develop a network model that captures contextual relationships between named entities for improved text characterization.
- To evaluate the model’s effectiveness in identifying unknown references in unstructured text.
Proposed method
- Constructs a network where nodes represent named entities (e.g., characters, locations, organizations) and edges are formed when entities co-occur within the same context.
- Uses computational simulations on book corpora to analyze topological properties such as shortest path length, clustering coefficient, and modularity.
- Applies the model to real-world text analysis tasks, particularly pattern recognition involving unknown reference resolution.
- Compares the proposed model’s performance against traditional word adjacency networks in identifying entities with unknown references.
- Employs topological metrics to assess structural complexity and organization of entity networks.
- Integrates the model as a complementary tool to statistical and deep learning approaches for text characterization.
Experimental results
Research questions
- RQ1How do named entity interactions form structured networks in written texts?
- RQ2What topological features emerge from modeling co-occurrence of named entities in textual contexts?
- RQ3How does the proposed model compare to word adjacency networks in identifying unknown references?
- RQ4To what extent does the network structure of named entities reflect underlying text organization?
- RQ5Can the model enhance the characterization of unstructured documents when combined with traditional methods?
Key findings
- The proposed model generates networks with short typical shortest path lengths, indicating efficient connectivity across entities.
- High clustering coefficients suggest strong local cohesion among named entities, reflecting thematic or narrative groupings.
- The networks exhibit modular organization, indicating distinct clusters of related entities, likely corresponding to narrative subplots or thematic sections.
- The model outperforms traditional word adjacency networks in identifying unknown references in text, demonstrating improved pattern recognition capability.
- The topological features—short paths, high clustering, and modularity—reveal complex, non-random organization in entity relationships.
- The model serves as a complementary approach to statistical and deep learning models, enhancing text characterization through structural analysis of named entities.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.