[论文解读] Simultaneous Speech Translation for Live Subtitling: from Delay to Display
本文提出了一种用于同时语音翻译(SimulST)的滚动行显示模式,通过预测字幕断点并以动态逐行滚动格式呈现翻译内容,从而在实时字幕中提升可读性。在英语→意大利语、德语和法语的实验中,滚动行模式实现了舒适的阅读速度(14.4–19.9 cps),85%–85%的字幕符合21 cps的舒适阅读阈值,同时将延迟保持在接近4秒的目标水平,优于逐词显示和基于块的显示模式。
With the increased audiovisualisation of communication, the need for live subtitles in multilingual events is more relevant than ever. In an attempt to automatise the process, we aim at exploring the feasibility of simultaneous speech translation (SimulST) for live subtitling. However, the word-for-word rate of generation of SimulST systems is not optimal for displaying the subtitles in a comprehensible and readable way. In this work, we adapt SimulST systems to predict subtitle breaks along with the translation. We then propose a display mode that exploits the predicted break structure by presenting the subtitles in scrolling lines. We compare our proposed mode with a display 1) word-for-word and 2) in blocks, in terms of reading speed and delay. Experiments on three language pairs (en$\ ightarrow$it, de, fr) show that scrolling lines is the only mode achieving an acceptable reading speed while keeping delay close to a 4-second threshold. We argue that simultaneous translation for readable live subtitles still faces challenges, the main one being poor translation quality, and propose directions for steering future research.
研究动机与目标
- 探究使用同时语音翻译(SimulST)进行实时多语言字幕生成的可行性。
- 解决SimulST中逐词显示模式可读性差的问题,该模式导致阅读速度波动且易造成观众疲劳。
- 通过预测的字幕断点,探索替代性显示策略——基于块的显示与滚动行显示。
- 评估不同显示模式在实时字幕中阅读速度与延迟之间的权衡。
- 识别将SimulST应用于实时字幕时的关键挑战,尤其是翻译质量与评估方法论问题。
提出的方法
- 作者微调一个SimulST模型,通过在生成过程中引入表示换行的特殊标记,使其能够预测字幕断点。
- 系统以流式方式逐字生成翻译,并在适当位置插入预测的断点。
- 实现了一种滚动行显示模式,实时逐行渲染字幕,同时保持语义块的完整性并支持动态滚动。
- 该方法对比了三种显示模式:逐词显示(逐字呈现)、基于块的显示(整句话一次性显示)和滚动行显示(动态逐行渲染)。
- 使用MuST-Cinema amara数据集,针对三种语言对(en→it、en→de、en→fr)测量阅读速度(rs)与延迟。
- 系统评估字幕是否符合21 cps的阅读速度阈值,并计算平均延迟(毫秒)以评估可用性。
实验结果
研究问题
- RQ1自动同时语音翻译能否成为生成实时多语言字幕的可行方法?
- RQ2SimulST系统生成模式在实时字幕可读性方面面临哪些挑战?
- RQ3不同显示策略——逐词、基于块和滚动行——如何影响实时字幕的阅读速度与延迟?
- RQ4预测的字幕断点能否提升SimulST生成字幕在实时场景下的可读性?
- RQ5在自动化实时字幕中,如何实现低延迟与可接受阅读速度之间的最佳平衡?
主要发现
- 滚动行模式在所有语言对中均实现了最低的平均阅读速度(14.4–19.9 cps),且85%–85%的字幕符合21 cps的舒适阅读阈值。
- 逐词显示模式导致最高的阅读速度(高达58.4 cps),但因单词停留时间不一致,可读性差且波动大。
- 基于块的显示模式延迟最高(4,503–5,273 ms),超过4秒阈值,且仅对高延迟系统略有改善阅读速度。
- 与基于块的显示相比,滚动行模式平均将延迟降低0.6秒,同时将延迟保持在接近4秒目标范围(4,090–4,708 ms)。
- 阅读速度与延迟之间的负相关关系得到验证,证实了这两个指标之间存在反比关系。
- en→de语言对在等待时间更长时(wait-5 vs. wait-3)阅读速度上升,但长度符合率下降(91% vs. 94%),凸显了准确断点预测的重要性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。