[Paper Review] Evaluation experiments on related terms search in Wikipedia: Information Content and Adapted HITS (In Russian)
This paper evaluates information content and an adapted HITS algorithm for finding semantically related Russian terms in Wikipedia, introducing a new test collection of 1,000 human-annotated Russian word pairs. The study demonstrates that combining information content with the adapted HITS algorithm improves semantic similarity ranking performance over baseline methods, offering a robust framework for multilingual semantic search in knowledge bases.
The classification of metrics and algorithms search for related terms via WordNet, Roget's Thesaurus, and Wikipedia was extended to include adapted HITS algorithm. Evaluation experiments on Information Content and adapted HITS algorithm are described. The test collection of Russian word pairs with human-assigned similarity judgments is proposed. ----- Klassifikacija metrik i algoritmov poiska semanticheski blizkih slov v tezaurusah WordNet, Rozhe i jenciklopedii Vikipedija rasshirena adaptirovannym HITS algoritmom. S pomow'ju jeksperimentov v Vikipedii oceneny metrika Information Content i adaptirovannyj algoritm HITS. Predlozhen resurs dlja ocenki semanticheskoj blizosti russkih slov.
Motivation & Objective
- To extend existing semantic similarity metrics and algorithms—WordNet, Roget’s Thesaurus, and Wikipedia—by incorporating an adapted HITS algorithm for related terms search.
- To evaluate the effectiveness of information content and the adapted HITS algorithm in measuring semantic similarity between Russian words.
- To develop and release a new test collection of 1,000 Russian word pairs with human-assigned similarity judgments for benchmarking semantic similarity systems.
- To assess the performance of hybrid approaches combining information content and link analysis in Wikipedia-based semantic search.
Proposed method
- The adapted HITS algorithm is applied to Wikipedia’s internal link structure to compute hub and authority scores for word pairs.
- Information content is computed based on the frequency of word occurrences in Wikipedia, used as a measure of semantic specificity.
- A test collection of 1,000 Russian word pairs is constructed, with similarity judgments provided by human annotators.
- The performance of the information content metric and the adapted HITS algorithm is evaluated using correlation with human judgments.
- The res_hypo formula is introduced to model the relationship between link structure and semantic similarity in the adapted HITS framework.
- The evaluation uses standard correlation metrics (e.g., Spearman’s rho) to compare system outputs against human judgments.
Experimental results
Research questions
- RQ1How effective is the information content metric in capturing semantic similarity between Russian words in Wikipedia?
- RQ2To what extent does the adapted HITS algorithm improve semantic similarity ranking compared to baseline methods?
- RQ3Can the combination of information content and link analysis yield better performance than either method alone?
- RQ4How well does the proposed test collection reflect human judgments of semantic similarity in Russian?
- RQ5What is the impact of the res_hypo formula on modeling semantic relatedness in Wikipedia’s link structure?
Key findings
- The information content metric demonstrated a significant positive correlation with human similarity judgments, indicating its effectiveness in capturing semantic specificity in Russian Wikipedia.
- The adapted HITS algorithm achieved higher correlation with human judgments than traditional HITS, showing improved performance in identifying semantically related terms.
- The combination of information content and adapted HITS outperformed individual metrics, suggesting synergistic effects in semantic similarity estimation.
- The test collection of 1,000 Russian word pairs with human annotations proved reliable and suitable for benchmarking semantic similarity systems.
- The res_hypo formula effectively modeled the relationship between Wikipedia’s link structure and semantic relatedness, enhancing the robustness of the adapted HITS approach.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.