[Paper Review] Evaluation of YTEX and MetaMap for clinical concept recognition
This study evaluates YTEX and MetaMap for clinical concept recognition in the 2013 ShARe/CLEF eHealth Task, using both systems as-is with post-processing to improve precision and recall. YTEX outperformed MetaMap in the relaxed task (F-score 4.6% higher) and showed 1.3% higher UMLS CUI mapping accuracy, while MetaMap achieved better precision in the strict task due to superior boundary detection after rule-based refinement.
We used MetaMap and YTEX as a basis for the construc- tion of two separate systems to participate in the 2013 ShARe/CLEF eHealth Task 1[9], the recognition of clinical concepts. No modifications were directly made to these systems, but output concepts were filtered using stop concepts, stop concept text and UMLS semantic type. Con- cept boundaries were also adjusted using a small collection of rules to increase precision on the strict task. Overall MetaMap had better per- formance than YTEX on the strict task, primarily due to a 20% perfor- mance improvement in precision. In the relaxed task YTEX had better performance in both precision and recall giving it an overall F-Score 4.6% higher than MetaMap on the test data. Our results also indicated a 1.3% higher accuracy for YTEX in UMLS CUI mapping.
Motivation & Objective
- To evaluate off-the-shelf clinical concept recognition tools, YTEX and MetaMap, in the context of the 2013 ShARe/CLEF eHealth Task.
- To assess the relative strengths and weaknesses of MetaMap’s heuristic-based disambiguation versus YTEX’s Lesk-based semantic similarity approach in clinical text.
- To improve concept recognition performance through post-processing rules that expand concept boundaries and filter low-value concepts.
- To analyze the impact of annotation guidelines and data characteristics on system evaluation outcomes.
- To determine which system offers better performance for clinical information extraction in real-world settings.
Proposed method
- YTEX 0.8 and MetaMap 2012 were used as-is, with no direct modifications to their core algorithms.
- A UIMA-based framework processed outputs from both systems, applying identical filtering and post-processing steps.
- Stop concepts (e.g., 'Disease', 'Injury') and stop concept text (e.g., 'mouse', 'mice') were filtered to reduce false positives.
- Concept boundaries were extended using a rule-based post-processing step that included preceding modifiers like 'LA', 'abd', 'chronic', and 'lower'.
- Only concepts matching required UMLS semantic types were retained, and results were evaluated on both training and test data.
- The 2012AB UMLS release was used for CUI mapping, and both systems were run with default parameters.
Experimental results
Research questions
- RQ1How do YTEX and MetaMap compare in precision, recall, and F-score on the strict and relaxed tasks of the ShARe/CLEF eHealth Task?
- RQ2What is the impact of rule-based boundary extension on concept recognition performance, particularly for modifiers and abbreviations?
- RQ3How does the choice of concept mapping strategy (heuristic-based vs. distributional similarity) affect accuracy and recall in clinical text?
- RQ4Why do performance differences emerge between the PTL dataset and the ShARe/CLEF test set, despite similar content?
- RQ5To what extent can post-processing rules improve system performance without modifying the underlying models?
Key findings
- MetaMap achieved a 20% improvement in precision on the strict task (0.722) compared to YTEX (0.512), leading to a higher F-score of 0.562 versus 0.473.
- YTEX outperformed MetaMap on the relaxed task, achieving a 4.6% higher F-score (0.924 vs. 0.878) due to better handling of partially mapped annotations.
- YTEX demonstrated 1.3% higher accuracy in UMLS CUI mapping (55.72% precision on PTL data) compared to MetaMap (80.28% precision, but lower recall).
- The PTL dataset showed significantly worse performance than the ShARe/CLEF test set, primarily due to stricter annotation guidelines penalizing broader concept mappings from YTEX.
- Abbreviations such as 'LA', 'abd', and 'MCA' were consistently misrecognized, and rule-based boundary extension significantly reduced this error class.
- Despite its lower performance on the strict task, YTEX showed superior ability to identify polysemous terms like 'inability' and 'call' in context, contributing to its relaxed task advantage.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.