[论文解读] Semantics-Aware Next-best-view Planning for Efficient Search and Detection of Task-relevant Plant Parts
本文提出了一种语义感知的下一最佳视角(NBV)规划策略,用于番茄温室中的机器人感知,通过语义类别标签和注意力机制优先检测与任务相关的植物部位(番茄、果柄、叶柄)。在仿真中,该方法实现了85.5%的检测准确率——比基线策略每株多检测4至11个部位——同时对遮挡、位置不确定性及植物复杂性保持鲁棒性。
Searching and detecting the task-relevant parts of plants is important to automate harvesting and de-leafing of tomato plants using robots. This is challenging due to high levels of occlusion in tomato plants. Active vision is a promising approach in which the robot strategically plans its camera viewpoints to overcome occlusion and improve perception accuracy. However, current active-vision algorithms cannot differentiate between relevant and irrelevant plant parts and spend time on perceiving irrelevant plant parts. This work proposed a semantics-aware active-vision strategy that uses semantic information to identify the relevant plant parts and prioritise them during view planning. The proposed strategy was evaluated on the task of searching and detecting the relevant plant parts using simulation and real-world experiments. In simulation experiments, the semantics-aware strategy proposed could search and detect 81.8% of the relevant plant parts using nine viewpoints. It was significantly faster and detected more plant parts than predefined, random, and volumetric active-vision strategies that do not use semantic information. The strategy proposed was also robust to uncertainty in plant and plant-part positions, plant complexity, and different viewpoint-sampling strategies. In real-world experiments, the semantics-aware strategy could search and detect 82.7% of the relevant plant parts using seven viewpoints, under complex greenhouse conditions with natural variation and occlusion, natural illumination, sensor noise, and uncertainty in camera poses. The results of this work clearly indicate the advantage of using semantics-aware active vision for targeted perception of plant parts and its applicability in the real world. It can significantly improve the efficiency of automated harvesting and de-leafing in tomato crop production.
研究动机与目标
- 解决在番茄温室中自动化采摘和去叶过程中,任务相关植物部位遮挡带来的机器人感知挑战。
- 克服传统主动视觉方法将所有植物部位同等对待、无法优先关注感兴趣对象(OOI)的局限。
- 通过将语义信息(类别标签和置信度分数)整合到下一最佳视角规划中,提升检测效率与准确性。
- 开发一种在线、自适应的注意力机制,动态引导视角选择,聚焦于OOI。
- 在类似真实环境的条件下,验证对植物结构不确定性、植物部位位置不确定性以及视角采样策略不确定性的鲁棒性。
提出的方法
- 将神经网络输出的语义分割结果(类别标签和置信度分数)集成到下一最佳视角(NBV)规划框架中。
- 使用注意力机制,在视角选择过程中为包含OOI(番茄、果柄、叶柄)的区域分配更高优先级。
- 采用体素表示(OctoMap)融合多视角观测结果,以处理三维植物结构中的不确定性。
- 提出一种新型语义NBV规划器,选择能最大化关于OOI的新信息的视角,而非仅关注整体场景覆盖。
- 在包含不同复杂度三维番茄植株模型的仿真环境中评估规划器性能,确保在受控且可重复的条件下进行测试。
- 采用50%的F1分数阈值判断对象检测的完整性,规划器在超过该阈值后仍持续感知对象,以确保完整性。

实验结果
研究问题
- RQ1将语义信息整合到主动视觉规划中,是否能提升在遮挡温室环境中对任务相关植物部位的检测效率与准确性?
- RQ2与传统的体素化、预设视角、随机视角及非语义主动视觉策略相比,语义感知NBV规划器在检测性能上表现如何?
- RQ3所提出的语义NBV规划器在植物部位位置不确定性、植物结构复杂性及视角采样约束方面,其鲁棒性达到何种程度?
- RQ4使用聚焦于OOI的注意力机制,是否能实现比非语义规划更快、更可靠的感兴趣植物部位检测?
- RQ5在不同对象检测完整性阈值下,该规划器表现如何?是否可在达到期望完整性后自动停止?
主要发现
- 在96次实验中,语义NBV规划器检测到了85.5%的所有植物部位,显著优于体素化NBV规划器,后者平均每株少检测4个部位。
- 与两种预设视角策略相比,语义NBV规划器分别多检测了5个和9个植物部位/株。
- 相比随机视角策略,语义NBV规划器每株多检测11个植物部位,证明其在目标感知方面的优越性。
- 在96次实验中,规划器的中位检测准确率达到88.9%,表明其在多种条件下均具有高度可靠性。
- 规划器对植物部位位置不确定性、植物复杂性变化以及不同视角采样策略均保持鲁棒,证实其在真实场景中的适用性。
- 由于规划器采用多视角机制,对象检测中的假阴性影响有限,即使某一视角失败,仍能从多个角度完成检测。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。