Skip to main content
QUICK REVIEW

[Paper Review] Language Tasks and Language Games: On Methodology in Current Natural Language Processing Research

David Schlangen|arXiv (Cornell University)|Aug 28, 2019
Topic Modeling40 references20 citations
TL;DR

This paper critiques the current methodology in natural language processing research by distinguishing between language tasks, micro-worlds, and dialogue games, arguing that progress toward modeling general language competence requires explicit grounding in cognitive capabilities. It advocates for clearer theoretical framing of new tasks and datasets using insights from linguistics and cognitive psychology to ensure they advance fundamental language understanding beyond superficial performance gains.

ABSTRACT

"This paper introduces a new task and a new dataset", "we improve the state of the art in X by Y" -- it is rare to find a current natural language processing paper (or AI paper more generally) that does not contain such statements. What is mostly left implicit, however, is the assumption that this necessarily constitutes progress, and what it constitutes progress towards. Here, we make more precise the normally impressionistically used notions of language task and language game and ask how a research programme built on these might make progress towards the goal of modelling general language competence.

Motivation & Objective

  • To clarify the implicit assumptions underlying the introduction of new language tasks and datasets in NLP research.
  • To distinguish and formalize the concepts of language tasks, micro-worlds, and dialogue games as distinct research objects.
  • To argue that progress in NLP should be evaluated not just by performance metrics, but by their contribution to modeling core language capabilities.
  • To advocate for stronger integration with linguistics and cognitive psychology to make theoretical claims about language competence explicit and contestable.
  • To provide a framework for evaluating new tasks and datasets based on their relevance to specific cognitive capabilities, such as compositionality or reference resolution.

Proposed method

  • Defining a language task as a mapping between input and output spaces, at least one of which involves natural language, with both intensional (e.g., 'translation') and extensional (dataset of input-output pairs) specifications.
  • Introducing 'micro-worlds' as simulated environments that generate disinterested responses to actions, enabling the study of situated, repeated language use.
  • Formalizing 'dialogue games' as repeated, structured interactions involving language tasks, regulated by rules and involving observable, responsive environments.
  • Proposing that models should be evaluated not only on task performance but on whether they internalize theoretically meaningful constructs, such as syntactic structure or compositional reasoning.
  • Recommending the use of 'data sheets' and explicit theoretical justification to ground new datasets and tasks in cognitive capabilities.
  • Using historical examples—such as visual question answering and the shapes dataset—to illustrate how task design can be improved by focusing on specific capabilities like compositionality.

Experimental results

Research questions

  • RQ1What distinguishes a language task from a dialogue game or a micro-world in NLP research?
  • RQ2How can the introduction of new datasets and tasks be justified if they do not advance the modeling of core language capabilities?
  • RQ3In what way do current evaluation practices obscure the true progress toward general language competence?
  • RQ4Why is it necessary to ground new NLP tasks and datasets in theoretical constructs from linguistics and cognitive psychology?
  • RQ5How can models be evaluated not just on performance, but on their internal representational structure and alignment with cognitive mechanisms?

Key findings

  • The paper identifies a recurring methodological flaw in NLP research: the assumption that introducing new tasks or datasets inherently constitutes progress, without explicit justification in terms of underlying language capabilities.
  • It demonstrates that many widely used datasets, such as those for visual question answering, suffer from 'language bias'—where models succeed without visual input—highlighting the need for more representative data.
  • The introduction of the shapes dataset by Andreas et al. (2016) exemplifies a better approach, as it explicitly targets the capability of compositional reasoning through programmatically generated questions with complex spatial relations.
  • The paper shows that synthetic datasets can be more valid for probing specific cognitive capabilities than natural-image-based datasets, which may introduce unintended biases.
  • It argues that models trained on tasks should be analyzed for internal structure (e.g., whether they form syntax trees), as this provides evidence for or against theoretical constructs like compositionality.
  • The paper concludes that stronger theoretical grounding—using constructs from linguistics and cognitive psychology—is essential to make claims about language competence explicit, contestable, and cumulative.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.