Skip to main content
QUICK REVIEW

[Paper Review] A Corpus of Sentence-level Revisions in Academic Writing: A Step towards Understanding Statement Strength in Communication

Chenhao Tan, Lillian Lee|arXiv (Cornell University)|May 6, 2014
Natural Language Processing Techniques22 references3 citations
TL;DR

This paper introduces the first large-scale corpus of sentence-level revisions in academic writing to study statement strength variations. By collecting and annotating 386 revision pairs from arXiv papers, the authors identify that 74.4% involve strength changes, with human annotators showing fair agreement (Fleiss’ Kappa = 0.322), revealing key insights into how linguistic features and domain-specific terms influence perceived strength in scientific communication.

ABSTRACT

The strength with which a statement is made can have a significant impact on the audience. For example, international relations can be strained by how the media in one country describes an event in another; and papers can be rejected because they overstate or understate their findings. It is thus important to understand the effects of statement strength. A first step is to be able to distinguish between strong and weak statements. However, even this problem is understudied, partly due to a lack of data. Since strength is inherently relative, revisions of texts that make claims are a natural source of data on strength differences. In this paper, we introduce a corpus of sentence-level revisions from academic writing. We also describe insights gained from our annotation efforts for this task.

Motivation & Objective

  • To address the understudied problem of statement strength in scientific communication, which affects peer review outcomes and public perception.
  • To create a large-scale, publicly available corpus of sentence-level revisions from academic papers to study strength differences.
  • To investigate how non-expert annotators perceive strength changes in technical statements, revealing discrepancies between expert and public interpretation.
  • To explore the impact of linguistic features—such as hedging, specificity, and domain terms—on perceived statement strength.
  • To enable future research on automated detection of strength changes and improved science communication.

Proposed method

  • Collected 108,000 sentence-level revision pairs from multi-version arXiv papers, filtering for valid, non-trivial changes.
  • Selected 1,000 random pairs after removing processing errors, then used Amazon Mechanical Turk for labeling with four labels: Stronger, Weaker, No Strength Change, I can’t tell.
  • Applied Fleiss’ Kappa to measure inter-annotator agreement, focusing only on pairs with absolute majority labels (n=386).
  • Analyzed annotator comments to identify patterns in reasoning, such as overreliance on specificity or domain-specific term perception.
  • Used a conservative labeling threshold to ensure reliability, discarding pairs without consensus.
  • Conducted qualitative analysis to uncover biases, such as length or specificity influencing perceived strength, even when irrelevant to logical strength.

Experimental results

Research questions

  • RQ1How do non-expert annotators perceive strength differences in technical statements from academic writing?
  • RQ2To what extent do linguistic features like hedging, specificity, and domain-specific terminology influence perceived statement strength?
  • RQ3What is the level of agreement among non-expert annotators in identifying strength changes in scientific revisions?
  • RQ4How do constraints and scope changes affect the perception of statement strength, especially when logical implications conflict with intuitive judgments?
  • RQ5Can annotator comments reveal generalizable features for modeling statement strength in scientific communication?

Key findings

  • Among 386 revision pairs with majority consensus, 74.4% were classified as strength changes, indicating that statement strength is frequently adjusted during academic writing.
  • Fleiss’ Kappa for the 386 pairs was 0.322, indicating fair agreement among annotators, despite the complexity of the task.
  • Annotators frequently labeled more specific statements as stronger, even when specificity did not affect logical strength, suggesting a heuristic bias toward detail.
  • Participants often interpreted domain-specific terms—such as 'vectors' vs. 'images' or 'adapt' vs. 'discover'—differently from experts, indicating a gap in public understanding.
  • Lengthier statements were often perceived as stronger, even when the added content was logically irrelevant, highlighting a perceptual bias in strength evaluation.
  • Annotator comments revealed that features like 'compelling evidence' were seen as stronger than 'compelling experimental evidence', suggesting scope reduction can be misinterpreted as strength increase.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.