Skip to main content
QUICK REVIEW

[论文解读] Live Target Detection with Deep Learning Neural Network and Unmanned Aerial Vehicle on Android Mobile Device

Ali Canberk Anar, Erkan Bostancı|arXiv (Cornell University)|Mar 19, 2018
Robotics and Sensor-Based Localization参考文献 2被引用 3
一句话总结

本文提出了一种基于Android的实时目标检测系统,采用DJI Phantom 3 Professional无人机和运行在移动硬件上的TensorFlow驱动的深度学习模型。系统实现了851毫秒的实时推理,分类置信度较高,例如在停车场场景中对'sports car'的接近度达到31.75%,展示了在移动平台上实现无人机视觉处理的可行性。

ABSTRACT

This paper describes the stages faced during the development of an Android program which obtains and decodes live images from DJI Phantom 3 Professional Drone and implements certain features of the TensorFlow Android Camera Demo application. Test runs were made and outputs of the application were noted. A lake was classified as seashore, breakwater and pier with the proximities of 24.44%, 21.16% and 12.96% respectfully. The joystick of the UAV controller and laptop keyboard was classified with the proximities of 19.10% and 13.96% respectfully. The laptop monitor was classified as screen, monitor and television with the proximities of 18.77%, 14.76% and 14.00% respectfully. The computer used during the development of this study was classified as notebook and laptop with the proximities of 20.04% and 11.68% respectfully. A tractor parked at a parking lot was classified with the proximity of 12.88%. A group of cars in the same parking lot were classified as sports car, racer and convertible with the proximities of 31.75%, 18.64% and 13.45% respectfully at an inference time of 851ms.

研究动机与目标

  • 开发一种实时、设备端的目标检测系统,用于在Android移动平台上处理无人机的实时视频流。
  • 将TensorFlow的移动推理引擎与DJI Phantom 3 Professional无人机的实时视频流集成。
  • 评估在不同环境条件下,设备端深度学习推理在分类现实世界物体方面的性能。
  • 展示在移动设备上部署轻量级CNN模型用于基于无人机的视觉感知的可行性。

提出的方法

  • 系统通过USB连接从DJI Phantom 3 Professional无人机捕获实时视频。
  • 原始视频帧被解码并预处理,以输入到预训练的TensorFlow MobileNet SSD模型中。
  • 推理引擎使用TensorFlow Lite框架在Android设备上运行,实现设备端处理。
  • 生成类别预测,并报告前五名类别的概率,以接近度分数形式表示与真实标签的接近程度。
  • 系统基于TensorFlow Android Camera Demo框架,实现摄像头集成与实时处理。
  • 对各种现实世界目标(包括车辆、电子设备和自然结构)的结果进行记录与分析。

实验结果

研究问题

  • RQ1轻量级深度学习模型是否能在使用商用无人机实时视频的Android移动设备上实现实时目标检测?
  • RQ2在分类多样化的现实世界目标(如车辆、电子设备和结构)时,设备端推理结果的准确性如何?
  • RQ3在消费级Android硬件上处理实时无人机视频时,移动优化的SSD模型的推理延迟是多少?
  • RQ4在真实世界条件下,不同物体类别之间的接近度分数(置信度水平)如何变化?
  • RQ5移动平台在多大程度上可以替代桌面级处理,用于基于无人机的视觉检测任务?

主要发现

  • 系统在移动Android设备上实现了平均851毫秒延迟的实时推理,用于目标检测。
  • 停车场中的一组汽车被分类为'sports car',接近度得分为31.75%,表明模型具有较高的置信度。
  • 拖拉机被分类为12.88%的接近度,反映出检测结果中等程度的置信度。
  • 笔记本电脑显示器被分类为'screen'、'monitor'和'television',接近度分别为18.77%、14.76%和14.00%。
  • 无人机控制器的操纵杆被分类为19.10%的接近度,表明对小型非标准物体的检测性能合理。
  • 开发过程中使用的电脑被分类为'notebook',接近度为20.04%,表明对便携式计算设备的有效识别。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。