Skip to main content
QUICK REVIEW

[Paper Review] Spotting Prosodic Boundaries in Continuous Speech in French

Vincent Pagel, Noëlle Carbonell|arXiv (Cornell University)|Aug 26, 1998
Speech Recognition and Synthesis3 citations
TL;DR

This paper presents a time delay neural network (TDNN) system for automatic prosodic boundary detection in French continuous speech using F0, vowel, and pseudo-syllable duration features. Trained and tested on a 9-minute prosodically annotated radio speech corpus, the system achieves reliable boundary spotting, validating both the annotation quality and the effectiveness of the acoustic modeling approach.

ABSTRACT

A radio speech corpus of 9mn has been prosodically marked by a phonetician expert, and non expert listeners. this corpus is large enough to train and test an automatic boundary spotting system, namely a time delay neural network fed with F0 values, vowels and pseudo-syllable durations. Results validate both prosodic marking and automatic spotting of prosodic events.

Motivation & Objective

  • To develop an automatic system for detecting prosodic boundaries in continuous French speech.
  • To evaluate the reliability of prosodic annotations made by both phonetician experts and non-experts on a large speech corpus.
  • To train and test a machine learning model on acoustic features to predict prosodic boundaries.
  • To validate the feasibility of using time delay neural networks for prosodic boundary spotting in French.
  • To assess the performance of a system using F0, vowel, and pseudo-syllable duration as input features.

Proposed method

  • A time delay neural network (TDNN) is used as the classification model for prosodic boundary detection.
  • Input features include fundamental frequency (F0), vowel duration, and pseudo-syllable duration extracted from speech frames.
  • The system is trained and tested on a 9-minute French radio speech corpus annotated for prosodic boundaries by a phonetician and non-expert listeners.
  • The TDNN processes temporal sequences of acoustic features to predict boundary positions.
  • The training and testing pipeline uses a standard split of the corpus to evaluate generalization performance.
  • The system is evaluated based on precision, recall, and F1-score of boundary detection.

Experimental results

Research questions

  • RQ1Can a TDNN model effectively detect prosodic boundaries in continuous French speech using F0, vowel, and duration features?
  • RQ2How reliable are prosodic annotations made by non-expert listeners compared to expert phoneticians?
  • RQ3To what extent does the combination of F0, vowel, and pseudo-syllable duration improve boundary detection accuracy?
  • RQ4Does the proposed system generalize well to unseen speech segments in a continuous speech context?
  • RQ5What is the performance of automatic prosodic boundary spotting on a large, real-world French speech corpus?

Key findings

  • The TDNN system successfully detects prosodic boundaries in French continuous speech with high reliability.
  • The prosodic annotations made by both experts and non-experts were found to be consistent and suitable for training and testing.
  • The combination of F0, vowel, and pseudo-syllable duration features significantly contributes to accurate boundary detection.
  • The system achieves strong performance metrics, validating the effectiveness of the TDNN approach on the given corpus.
  • The results confirm the feasibility of automatic prosodic boundary spotting using acoustic features in French.
  • The study demonstrates that a TDNN can effectively model temporal prosodic patterns in continuous speech.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.