[Paper Review] A Comparison of WordNet and Roget's Taxonomy for Measuring Semantic Similarity
This paper evaluates Roget's International Thesaurus as a semantic taxonomy for measuring word similarity, applying four established metrics to assess semantic relatedness. It finds that the traditional edge-counting approach on Roget's taxonomy achieves a high correlation of r=0.88 with human judgments, approaching the human upper bound of r=0.90.
This paper presents the results of using Roget's International Thesaurus as the taxonomy in a semantic similarity measurement task. Four similarity metrics were taken from the literature and applied to Roget's The experimental evaluation suggests that the traditional edge counting approach does surprisingly well (a correlation of r=0.88 with a benchmark set of human similarity judgements, with an upper bound of r=0.90 for human subjects performing the same task.)
Motivation & Objective
- To evaluate Roget's International Thesaurus as an alternative to WordNet for semantic similarity measurement.
- To assess the performance of four existing similarity metrics when applied to Roget's taxonomy.
- To determine how well traditional edge-counting methods on Roget's taxonomy align with human semantic judgments.
- To establish a benchmark for Roget-based semantic similarity using human-annotated similarity judgments.
Proposed method
- Four semantic similarity metrics from the literature were applied to Roget's International Thesaurus as the underlying taxonomy.
- The taxonomy structure was used to compute path lengths between word pairs, forming the basis for edge-counting similarity.
- Similarity scores were computed based on the shortest path between concepts in Roget's hierarchical structure.
- A standard benchmark set of human similarity judgments was used for evaluation and correlation analysis.
- Pearson correlation coefficients were calculated to compare the system's similarity scores with human judgments.
Experimental results
Research questions
- RQ1How does Roget's taxonomy compare to WordNet in measuring semantic similarity?
- RQ2What is the performance of edge-counting methods when applied to Roget's taxonomy?
- RQ3How close does the edge-counting approach on Roget's taxonomy come to human performance on the same task?
- RQ4Can Roget's taxonomy serve as a viable alternative to WordNet for semantic similarity tasks?
Key findings
- The edge-counting method on Roget's taxonomy achieved a correlation of r=0.88 with human similarity judgments.
- This performance is remarkably close to the upper bound of r=0.90 observed in human subjects performing the same task.
- The results indicate that Roget's taxonomy is a strong candidate for semantic similarity measurement, rivaling established WordNet-based approaches.
- The traditional edge-counting approach performs surprisingly well on Roget's taxonomy, outperforming expectations given its simplicity.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.