[论文解读] ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on Signboard
本论文介绍了 ICDAR 2019 鲁棒阅读挑战赛中的中文路牌文本(ReCTS)任务,提出一个大规模数据集,包含 25,000 幅带详细字符与文本行标注的路牌图像。挑战赛包含四项任务——字符识别、文本行识别、文本行检测与端到端识别,采用一种新颖的多真实标注(multi-ground truth)评估方法,以应对中文文本的固有歧义,共吸引了来自全球机构的 46 支团队、262 份提交,表现优异。
Chinese scene text reading is one of the most challenging problems in computer vision and has attracted great interest. Different from English text, Chinese has more than 6000 commonly used characters and Chinesecharacters can be arranged in various layouts with numerous fonts. The Chinese signboards in street view are a good choice for Chinese scene text images since they have different backgrounds, fonts and layouts. We organized a competition called ICDAR2019-ReCTS, which mainly focuses on reading Chinese text on signboard. This report presents the final results of the competition. A large-scale dataset of 25,000 annotated signboard images, in which all the text lines and characters are annotated with locations and transcriptions, were released. Four tasks, namely character recognition, text line recognition, text line detection and end-to-end recognition were set up. Besides, considering the Chinese text ambiguity issue, we proposed a multi ground truth (multi-GT) evaluation method to make evaluation fairer. The competition started on March 1, 2019 and ended on April 30, 2019. 262 submissions from 46 teams are received. Most of the participants come from universities, research institutes, and tech companies in China. There are also some participants from the United States, Australia, Singapore, and Korea. 21 teams submit results for Task 1, 23 teams submit results for Task 2, 24 teams submit results for Task 3, and 13 teams submit results for Task 4. The official website for the competition is http://rrc.cvc.uab.es/?ch=12.
研究动机与目标
- 为应对真实路牌中中文文本所面临的复杂版式、多样字体与高字符变异性等挑战。
- 通过构建一个大规模、精细标注的 25,000 幅路牌图像数据集,建立鲁棒的中文场景文本识别基准。
- 在多个任务上评估并比较最先进方法的表现:字符识别、文本行识别、文本行检测与端到端识别。
- 提出一种多真实标注(multi-GT)评估策略,公平反映中文文本转录中的固有歧义。
- 通过组织全球性竞赛,促进大学、研究机构与科技公司之间的国际合作与创新。
提出的方法
- 收集并精心标注了 25,000 幅街景路牌图像,为每个字符与文本行提供边界框与转录结果。
- 定义了四项独立任务:字符识别、文本行识别、文本行检测与端到端识别,以评估系统在不同能力层级的表现。
- 提出一种多真实标注(multi-GT)评估方法,允许多个正确转录结果对应同一图像,以反映中文文本识别中的歧义性。
- 通过专用网站在线举办竞赛,共吸引来自 10 个国家(包括中国、美国、澳大利亚、新加坡与韩国)的 46 支团队参与。
- 设计评估指标以应对转录差异,确保在不同识别方法之间实现公平比较。
实验结果
研究问题
- RQ1如何构建一个大规模、高质量的中文路牌图像数据集,以支持鲁棒文本识别研究?
- RQ2由于版式、字体与字符复杂性,识别真实路牌中的中文文本面临哪些关键挑战?
- RQ3多真实标注评估方法如何提升中文文本识别系统评估的公平性与准确性?
- RQ4在真实世界中文路牌数据上,字符识别、文本行识别、检测与端到端识别等任务的性能水平如何?
- RQ5来自不同机构的最先进方法在真实环境下处理中文场景文本复杂性方面表现如何?
主要发现
- 竞赛共收到 262 份提交,来自 46 支团队,显示出全球在中文文本识别研究中的高度参与。
- 21 支团队提交了任务 1(字符识别)结果,23 支团队提交了任务 2(文本行识别)结果,24 支团队提交了任务 3(文本行检测)结果,13 支团队提交了任务 4(端到端识别)结果。
- 多真实标注评估方法有效减少了因转录歧义带来的偏差,实现了更公平、更可靠的性能比较。
- 发布的 25,000 幅标注路牌图像数据集,为未来中文场景文本理解研究提供了全面的基准。
- 所有任务均取得了高性能表现,尤其在端到端识别任务中表现突出,表明检测与识别统一建模方面取得进展。
- 挑战赛凸显了中国团队的主导地位,同时美国、澳大利亚、新加坡与韩国的机构也作出了显著贡献。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。