Skip to main content
QUICK REVIEW

[Paper Review] Defining maximum acceptable latency of AI-enhanced CAI tools

Claudio Fantinuoli, Maddalena Montecchio|arXiv (Cornell University)|Jan 8, 2022
Interpreting and Communication in Healthcare4 citations
TL;DR

This study investigates the maximum acceptable latency for AI-enhanced computer-assisted interpreting (CAI) tools in simultaneous interpreting. Through an empirical experiment with professional interpreters, it finds that system latency up to 3 seconds does not significantly impair accuracy or fluency, suggesting that higher-latency, context-aware NLP features can be safely integrated into next-generation CAI tools without compromising performance.

ABSTRACT

Recent years have seen an increasing number of studies around the design of computer-assisted interpreting tools with integrated automatic speech processing and their use by trainees and professional interpreters. This paper discusses the role of system latency of such tools and presents the results of an experiment designed to investigate the maximum system latency that is cognitively acceptable for interpreters working in the simultaneous modality. The results show that interpreters can cope with a system latency of 3 seconds without any major impact in the rendition of the original text, both in terms of accuracy and fluency. This value is above the typical latency of available AI-based CAI tools and paves the way to experiment with larger context-based language models and higher latencies.

Motivation & Objective

  • To determine the maximum system latency that interpreters can cognitively tolerate when using AI-enhanced CAI tools in simultaneous interpreting.
  • To assess how increasing latency affects the accuracy, fluency, and cognitive load in interpreter performance.
  • To evaluate whether current AI-based CAI tools with moderate latency can be safely extended to support more complex, context-based NLP features.
  • To empirically establish a threshold for system latency that preserves the quality of interpretation without disrupting the ear-voice span (EVS).

Proposed method

  • An experimental setup was designed in which professional interpreters interpreted live speech with artificially introduced delays in number suggestions from an ASR simulation system.
  • Latency was systematically varied across five conditions: 1, 2, 3, 4, and 5 seconds.
  • Performance was evaluated using both stimuli-based and segment-based assessments, focusing on number accuracy, referent accuracy, disfluency, and fluency.
  • Interpreters’ renditions were scored for precision, fluency, and paralinguistic quality across different latency levels.
  • The evaluation framework included both quantitative metrics (accuracy percentages) and qualitative ratings (fluency, disfluency).
  • The study used a within-subjects design with multiple interpreters to assess performance trends across latency conditions.

Experimental results

Research questions

  • RQ1What is the maximum system latency that interpreters can tolerate without a significant decline in interpretation accuracy?
  • RQ2How does increasing system latency affect the fluency and disfluency in interpreter renditions?
  • RQ3Does a latency of 3 seconds or less have a measurable impact on the cognitive load or quality of simultaneous interpretation?
  • RQ4Can interpreters successfully adapt their ear-voice span (EVS) to accommodate suggestions up to 3 seconds delay without compromising performance?

Key findings

  • The highest number accuracy (98.85%) was achieved at a 3-second latency, indicating possible adaptation or learning effects over time.
  • At 4 seconds latency, number accuracy dropped to 93.14%, and at 5 seconds, it further declined to 94.86%, showing a measurable negative impact.
  • Referent accuracy peaked at 100% with 3 seconds latency and dropped to 94.28% at 4 seconds and 85.71% at 5 seconds.
  • Disfluency in number pronunciation reached 25.71% at 4 seconds, the highest among all conditions, indicating increased cognitive strain.
  • Fluency ratings declined significantly at 4 seconds and were lowest at 5 seconds, indicating a deterioration in delivery quality.
  • Segment-level accuracy remained high up to 3 seconds latency, with a minimal decline observed from 4 seconds onward, suggesting a threshold effect at 3 seconds.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.