Skip to main content
QUICK REVIEW

[Paper Review] Shallow Discourse Annotation for Chinese TED Talks

Wanqiu Long, Xinyi Cai|arXiv (Cornell University)|Mar 9, 2020
Natural Language Processing Techniques15 references4 citations
TL;DR

This paper presents a novel, publicly available corpus of Chinese TED talks annotated with shallow discourse relations using an adapted PDTB-3 scheme, tailored to the features of Chinese spoken discourse. The annotation achieves high inter-annotator agreement (kappa > 0.85) and reveals a more balanced distribution of discourse relations—particularly higher proportions of Contingency, Temporal, and Comparison—compared to existing news-focused Chinese corpora.

ABSTRACT

Text corpora annotated with language-related properties are an important resource for the development of Language Technology. The current work contributes a new resource for Chinese Language Technology and for Chinese-English translation, in the form of a set of TED talks (some originally given in English, some in Chinese) that have been annotated with discourse relations in the style of the Penn Discourse TreeBank, adapted to properties of Chinese text that are not present in English. The resource is currently unique in annotating discourse-level properties of planned spoken monologues rather than of written text. An inter-annotator agreement study demonstrates that the annotation scheme is able to achieve highly reliable results.

Motivation & Objective

  • To develop a high-quality, publicly accessible discourse-annotated corpus of Chinese spoken monologues, specifically TED talks.
  • To adapt the PDTB-3 discourse annotation scheme to account for unique features of Chinese spoken discourse, such as implicit relations and connective usage.
  • To evaluate the reliability of the annotation scheme through inter-annotator agreement studies.
  • To compare the discourse relation distribution in spoken Chinese with that in written Chinese corpora, particularly news reports.
  • To provide a resource for training and evaluating discourse-aware NLP models in Chinese, especially for spoken language applications.

Proposed method

  • Adapted the PDTB-3 discourse annotation framework to Chinese spoken discourse, modifying sense hierarchies by removing AltLexC and adding Progression.
  • Defined detailed annotation guidelines and criteria for discourse relations, argument scope, and connective identification in spoken Chinese.
  • Conducted rigorous training and quality control for annotators, including iterative review and consensus-building processes.
  • Collected and annotated 3212 discourse relations across 32 TED talks (in Chinese or translated into Chinese), focusing on intra- and cross-sentence relations.
  • Performed inter-annotator agreement studies using Fleiss’ Kappa to validate reliability, achieving >0.85 for most metrics.
  • Conducted comparative analysis of discourse relation distributions between the new corpus and the CUHK-DTBC (a news-focused Chinese corpus).

Experimental results

Research questions

  • RQ1How can the PDTB-3 discourse annotation scheme be effectively adapted to Chinese spoken discourse, particularly to address features absent in English?
  • RQ2What is the inter-annotator agreement level when applying the adapted scheme to Chinese TED talks, indicating its reliability?
  • RQ3How does the distribution of discourse relations in Chinese spoken TED talks differ from that in written Chinese corpora like CUHK-DTBC?
  • RQ4To what extent do explicit and implicit discourse relations co-occur in spoken Chinese, and how does this compare to written Chinese?
  • RQ5What are the most frequent discourse sense types in Chinese spoken monologues, and how do they reflect the rhetorical structure of spoken discourse?

Key findings

  • The adapted annotation scheme achieved high inter-annotator agreement, with Fleiss’ Kappa values exceeding 0.85 for most annotation indicators, confirming its reliability.
  • The corpus contains 3212 annotated discourse relations, with approximately equal proportions of explicit and implicit relations, suggesting a higher prevalence of explicit connectives in spoken Chinese compared to written text.
  • The distribution of discourse relation types in the corpus is more balanced than in news-focused corpora: Contingency (27.6%), Temporal (19.2%), and Comparison (15.7%) are more frequent than Expansion (37.5%), which contrasts sharply with the dominance of Expansion (52%) in the CUHK-DTBC.
  • The most frequent second-level discourse senses are Cause (20%), Conjunction (13%), and Concession (13%), with the top 10 senses accounting for 86% of all annotations.
  • The corpus reveals distinct discourse patterns in Chinese spoken discourse, including a higher frequency of Contingency and Comparison relations, indicating differences from written text genres.
  • The findings demonstrate that spoken Chinese TED talks exhibit a more diverse and less expansion-dominated discourse structure than written Chinese news reports, supporting the value of genre-specific corpora for NLP.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.