[论文解读] Testing of Autonomous Driving Systems: Where Are We and Where Should We Go?
本文通过访谈10名自动驾驶系统(ADS)开发者及对100名从业者进行调查,开展了一项关于自动驾驶系统(ADS)测试实践与需求的全面实证研究。研究识别出七种常见的测试实践和四种新兴需求——边缘案例检测、测试加速、场景构建以及数据标注——同时揭示了现有软件工程(SE)研究中的显著空白,即大多忽视多模块系统,且对测试速度和工具支持等实际需求关注不足。
Autonomous driving has shown great potential to reform modern transportation. Yet its reliability and safety have drawn a lot of attention and concerns. Compared with traditional software systems, autonomous driving systems (ADSs) often use deep neural networks in tandem with logic-based modules. This new paradigm poses unique challenges for software testing. Despite the recent development of new ADS testing techniques, it is not clear to what extent those techniques have addressed the needs of ADS practitioners. To fill this gap, we present the first comprehensive study to identify the current practices and needs of ADS testing. We conducted semi-structured interviews with developers from 10 autonomous driving companies and surveyed 100 developers who have worked on autonomous driving systems. A systematic analysis of the interview and survey data revealed 7 common practices and 4 emerging needs of autonomous driving testing. Through a comprehensive literature review, we developed a taxonomy of existing ADS testing techniques and analyzed the gap between ADS research and practitioners' needs. Finally, we proposed several future directions for SE researchers, such as developing test reduction techniques to accelerate simulation-based ADS testing.
研究动机与目标
- 理解自动驾驶系统(ADS)在实际开发环境中的当前工业测试实践。
- 识别现有软件工程(SE)研究尚未充分满足的自动驾驶系统(ADS)从业者所面临的新兴测试需求。
- 评估现有SE研究在ADS测试方面与工业实践需求之间的契合度。
- 通过基于实证证据的未来研究方向,弥合学术研究与工业实践之间的差距。
提出的方法
- 对来自不同自动驾驶公司的10名开发者进行了半结构化访谈,以探索实际测试实践与挑战。
- 设计并发放了针对1,978名ADS开发者的大型调查,收集到100份有效回复,用于量化和验证访谈发现。
- 对117篇SE会议和期刊论文进行了系统性文献综述,筛选出42篇聚焦于ADS测试的论文,以构建现有技术的分类体系。
- 通过三角验证法整合定性访谈数据、定量调查结果与文献综述发现,以识别模式、空白点与研究机会。
- 采用主题分析法,并通过组间一致性检验(Cohen’s Kappa = 0.85)确保访谈与调查数据编码的一致性。
- 基于识别出的研究空白,提出了未来研究方向,尤其聚焦于测试缩减、场景生成以及复杂ADS工作流的工具支持。
实验结果
研究问题
- RQ1当前工业界在测试自动驾驶系统方面有哪些实践?
- RQ2ADS从业者在测试方面有哪些新兴需求,特别是在工具与方法论方面?
- RQ3现有软件工程研究在ADS测试方面的技术与工业界需求和实践之间的契合度如何?
- RQ4在实际应用中,端到端模型与多模块ADS架构的测试实践有何不同?
主要发现
- 识别出七种常见的测试实践,包括系统级指标(如一致性与延迟),这些指标通常未被标准模型准确率指标所涵盖。
- 揭示了四项关键新兴需求:识别边缘案例与意外驾驶场景、加速测试(尤其是仿真测试)、支持构建复杂驾驶场景的工具,以及数据标注支持。
- 尽管边缘案例检测在SE研究中已有充分覆盖,但测试加速与场景/工具支持在现有文献中仍研究不足。
- 大多数现有ADS测试技术针对的是端到端深度学习模型,但工业系统越来越多地依赖将感知网络与基于逻辑的控制器相结合的模块化架构。
- 当前SE研究与工业需求之间存在显著脱节,尤其体现在仿真测试的实用工具支持与性能优化方面。
- 本研究将测试缩减技术识别为未来关键研究方向,以加速仿真测试,而这是工业ADS开发中的主要瓶颈。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。