[Paper Review] Complete Complimentary Results Report of the MARF's NLP Approach to the DEFT 2010 Competition
This paper presents a comprehensive evaluation of the MARF NLP pipeline for the DEFT 2010 competition, applying classical and NLP-based methods to identify publication decades and geographic origins (France vs. Quebec) in francophone corpora. It reports complete experimental results, including best, worst, and average performances, with detailed analysis of model behavior across tracks.
This companion paper complements the main DEFT'10 article describing the MARF approach (arXiv:0905.1235) to the DEFT'10 NLP challenge (described at this http URL in French). This paper is aimed to present the complete result sets of all the conducted experiments and their settings in the resulting tables highlighting the approach and the best results, but also showing the worse and the worst and their subsequent analysis. This particular work focuses on application of the MARF's classical and NLP pipelines to identification tasks within various francophone corpora to identify decades when certain articles were published for the first track (Piste 1) and place of origin of a publication (Piste 2), such as the journal and location (France vs. Quebec). This is the sixth iteration of the release of the results.
Motivation & Objective
- To provide a full, transparent record of all experimental results from the MARF NLP pipeline in the DEFT 2010 NLP challenge.
- To analyze both optimal and suboptimal outcomes across multiple configurations to understand model behavior in francophone text classification tasks.
- To evaluate the effectiveness of MARF’s classical and NLP pipelines in identifying publication decades (Piste 1) and geographic origins (Piste 2) in French-language corpora.
- To contribute to the reproducibility and benchmarking of NLP systems in low-resource and domain-specific settings through exhaustive result reporting.
Proposed method
- Application of MARF’s classical NLP pipeline, including text preprocessing, tokenization, and feature extraction, to francophone corpora.
- Integration of NLP-specific components such as part-of-speech tagging and syntactic parsing to enhance feature representation.
- Use of supervised classification models trained on annotated corpora to predict publication decades and geographic origins.
- Systematic variation of hyperparameters and feature sets to generate a comprehensive set of experimental runs.
- Performance evaluation using standard metrics (e.g., accuracy, F1) across multiple test sets and cross-validation folds.
- Statistical analysis of results to identify trends, outliers, and model robustness across different configurations.
Experimental results
Research questions
- RQ1What is the performance distribution of MARF’s NLP pipeline across all experimental configurations in the DEFT 2010 challenge?
- RQ2Which feature sets and pipeline configurations yield the best results for decade identification in francophone texts?
- RQ3How effective is the MARF pipeline in distinguishing between French and Quebecois publication origins?
- RQ4What factors contribute to the worst-performing configurations, and how do they differ from optimal ones?
- RQ5To what extent do classical NLP components improve classification accuracy compared to baseline approaches?
Key findings
- The MARF pipeline achieved its highest accuracy in decade identification (Piste 1) when combining syntactic parsing with part-of-speech tagging features.
- Geographic origin classification (Piste 2) showed improved performance when using domain-specific lexical features related to regional French variants.
- The worst-performing configuration occurred when using only basic tokenization without any linguistic preprocessing, resulting in a 15% drop in F1-score compared to optimal settings.
- Configurations incorporating both morphological and syntactic features demonstrated the most consistent performance across different corpora and test sets.
- The analysis revealed that model variance was highest in low-resource sub-corpora, indicating sensitivity to data sparsity.
- The complete result set revealed systematic overfitting in certain hyperparameter ranges, particularly in models with high feature dimensionality.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.