Skip to main content
QUICK REVIEW

[Paper Review] Teach Me to Explain: A Review of Datasets for Explainable NLP.

Sarah Wiegreffe, Ana Marasović|arXiv (Cornell University)|Feb 24, 2021
Topic ModelingComputer Science143 references59 citations
TL;DR

This paper reviews datasets for explainable NLP, categorizing human-annotated explanations into three types—highlights, free-text, and structured—and synthesizes findings on their collection, use, and evaluation. It provides recommendations for future dataset creation based on lessons learned from existing literature and practices in data augmentation, model training, and explanation quality assessment.

ABSTRACT

Explainable NLP (ExNLP) has increasingly focused on collecting human-annotated explanations. These explanations are used downstream in three ways: as data augmentation to improve performance on a predictive task, as a loss signal to train models to produce explanations for their predictions, and as a means to evaluate the quality of model-generated explanations. In this review, we identify three predominant classes of explanations (highlights, free-text, and structured), organize the literature on annotating each type, point to what has been learned to date, and give recommendations for collecting ExNLP datasets in the future.

Motivation & Objective

  • To identify and categorize the three main types of human-annotated explanations used in explainable NLP: highlights, free-text, and structured explanations.
  • To organize and synthesize existing literature on the annotation of each explanation type, highlighting methodological trends and challenges.
  • To summarize key findings from current research on how explanations are used—specifically as data augmentation, loss signals, and evaluation metrics.
  • To provide actionable recommendations for future collection of high-quality ExNLP datasets based on empirical insights and best practices.
  • To support the development of more reliable, interpretable, and generalizable NLP models through improved dataset design and annotation standards.

Proposed method

  • Systematic review of existing datasets and annotation practices in explainable NLP, focusing on three explanation types: highlights, free-text, and structured.
  • Classification of datasets based on annotation format, task type, and downstream application (e.g., data augmentation, model training, evaluation).
  • Analysis of how explanations are used in three primary ways: improving model performance via data augmentation, training models to generate explanations via loss signals, and evaluating model explanations.
  • Synthesis of findings across studies to identify common challenges, design patterns, and best practices in explanation collection.
  • Development of recommendations for future dataset collection, emphasizing consistency, scalability, and alignment with model evaluation needs.
  • Use of qualitative and comparative analysis to assess the quality and utility of existing datasets in supporting ExNLP research.

Experimental results

Research questions

  • RQ1What are the dominant forms of human-annotated explanations in NLP, and how do they differ in structure and purpose?
  • RQ2How are explanations currently used in downstream NLP tasks, and what impact do they have on model performance and interpretability?
  • RQ3What methodological patterns and challenges emerge in the annotation of highlights, free-text, and structured explanations?
  • RQ4What lessons can be drawn from existing datasets to inform the design of future ExNLP datasets?
  • RQ5How can future datasets be optimized to support data augmentation, model training, and explanation quality evaluation?

Key findings

  • Highlights, free-text, and structured explanations represent the three primary categories of human-annotated explanations in ExNLP, each with distinct annotation practices and use cases.
  • Explanations are commonly used as data augmentation to improve model performance on predictive tasks, particularly in low-resource settings.
  • Using explanations as a loss signal during training helps align model-generated explanations with human-annotated ones, improving faithfulness and interpretability.
  • Evaluation of model-generated explanations is most effective when grounded in human-annotated explanations, which serve as a gold standard for comparison.
  • Despite progress, inconsistencies in annotation guidelines and evaluation protocols remain a challenge across datasets, limiting reproducibility and comparability.
  • Future datasets should prioritize standardized, scalable, and diverse annotation protocols to support robust model development and evaluation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.