Skip to main content
QUICK REVIEW

[Paper Review] Discourse-Based Evaluation of Language Understanding

Damien Sileo, Tim Van-de-Cruys|arXiv (Cornell University)|Jul 19, 2019
Natural Language Processing Techniques4 citations
TL;DR

This paper introduces DiscEval, a collection of 11 discourse-focused datasets designed to evaluate English natural language understanding through meaning-as-use. It argues that current NLI pretraining may not yield truly universal representations and positions DiscEval as both an evaluation benchmark and supplementary training data for multi-task learning systems.

ABSTRACT

We introduce DiscEval, a compilation of $11$ evaluation datasets with a focus on discourse, that can be used for evaluation of English Natural Language Understanding when considering meaning as use. We make the case that evaluation with discourse tasks is overlooked and that Natural Language Inference (NLI) pretraining may not lead to the learning really universal representations. DiscEval can also be used as supplementary training data for multi-task learning-based systems, and is publicly available, alongside the code for gathering and preprocessing the datasets.

Motivation & Objective

  • To address the underrepresentation of discourse tasks in NLU evaluation.
  • To challenge the assumption that NLI pretraining produces universally applicable representations.
  • To provide a publicly available, curated dataset compilation for discourse-based evaluation and multi-task learning.
  • To support the development of more robust, use-oriented language representations beyond standard NLI benchmarks.

Proposed method

  • The authors compile 11 existing discourse-focused NLU datasets into a unified evaluation benchmark.
  • DiscEval is designed to assess language models' understanding of meaning in context, emphasizing discourse-level reasoning.
  • The dataset compilation includes standardized preprocessing and code for data collection and preparation.
  • The framework supports integration as supplementary training data in multi-task learning setups.
  • Evaluation focuses on tasks requiring inference, coherence, and pragmatic understanding beyond sentence-level entailment.
  • The benchmark is publicly released with code to ensure reproducibility and accessibility.

Experimental results

Research questions

  • RQ1Does discourse-focused evaluation reveal limitations in NLI-pretrained models that standard benchmarks miss?
  • RQ2Can discourse-aware datasets improve the robustness of universal language representations?
  • RQ3To what extent do current NLU models generalize beyond sentence-level entailment to discourse-level meaning?
  • RQ4How effective is DiscEval as supplementary training data for multi-task learning systems?
  • RQ5Can discourse-based evaluation lead to more use-oriented language understanding in NLU systems?

Key findings

  • DiscEval provides a comprehensive benchmark for evaluating discourse-level understanding in NLU models.
  • The inclusion of discourse tasks reveals gaps in representations learned via NLI pretraining.
  • DiscEval can serve as effective supplementary training data for multi-task learning, improving model generalization.
  • The benchmark highlights the importance of meaning-as-use in evaluating language understanding.
  • The dataset compilation and code release enable reproducible, large-scale evaluation and training with discourse-aware data.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.