Skip to main content
QUICK REVIEW

[논문 리뷰] Comparison and Benchmarking of AI Models and Frameworks on Mobile Devices

Chunjie Luo, Xiwen He|arXiv (Cornell University)|2020. 05. 07.
Advanced Neural Network Applications참고 문헌 34인용 수 41
한 줄 요약

AIoTBench 벤치마크는 6개 모델과 3개의 프레임워크에서 모바일 AI 추론을 수행하고 5대 디바이스에서 평가하며, VIPS와 VOPS로 점수를 매겨 AI 능력과 트레이드오프를 비교한다.

ABSTRACT

Due to increasing amounts of data and compute resources, deep learning achieves many successes in various domains. The application of deep learning on the mobile and embedded devices is taken more and more attentions, benchmarking and ranking the AI abilities of mobile and embedded devices becomes an urgent problem to be solved. Considering the model diversity and framework diversity, we propose a benchmark suite, AIoTBench, which focuses on the evaluation of the inference abilities of mobile and embedded devices. AIoTBench covers three typical heavy-weight networks: ResNet50, InceptionV3, DenseNet121, as well as three light-weight networks: SqueezeNet, MobileNetV2, MnasNet. Each network is implemented by three frameworks which are designed for mobile and embedded devices: Tensorflow Lite, Caffe2, Pytorch Mobile. To compare and rank the AI capabilities of the devices, we propose two unified metrics as the AI scores: Valid Images Per Second (VIPS) and Valid FLOPs Per Second (VOPS). Currently, we have compared and ranked 5 mobile devices using our benchmark. This list will be extended and updated soon after.

연구 동기 및 목표

  • 모바일 및 임베디드 디바이스에서 온디바이스 AI 추론 벤치마크의 필요성과 동기를 제시한다.
  • 모델 아키텍처와 프레임워크 전반에 걸친 추론을 평가하기 위해 AIoTBench를 제안한다.
  • 디바이스 간 비교를 위한 통합 AI 점수 VIPS와 VOPS를 소개한다.
  • 실제 모바일 하드웨어에서 실용적인 벤치마킹 워크플로우를 제공한다.

제안 방법

  • 여섯 개 네트워크(세 가지 무거운: ResNet50, InceptionV3, DenseNet121; 세 가지 경량: SqueezeNet, MobileNetV2, MnasNet)로 워크로드 정의.
  • 각 모델을 세 가지 모바일/임베디드 프레임워크(TensorFlow Lite, Caffe2, PyTorch Mobile)로 구현.
  • 추론 벤치마킹을 위해 ImageNet 검증 세트(클래스당 5개의 임의 이미지, 총 5000개) 사용.
  • 프레임워크별 표준화된 정규화, 형태, 색상 순서를 적용하여 입력을 전처리한다(표 IV에 명시).
  • 정확도 및 이미지별 추론 시간을 측정하고, 정의된 식으로 VIPS 및 VOPS AI 점수를 계산한다.

실험 결과

연구 질문

  • RQ1다른 AI 모델이 모바일 기기에서 프레임워크 간 정확도, 크기 및 추론 속도 간의 트레이드오프를 어떻게 보이는가?
  • RQ2모바일 하드웨어와 프레임워크가 온-디바이스 AI 추론 성능에 어떤 영향을 미치는가?
  • RQ3통합 VIPS와 VOPS 점수가 모바일 디바이스를 비교하는 데 얼마나 안정적인 기반을 제공하는가?
  • RQ4동일한 디바이스에서 프레임워크 선택이 모델 성능에 어떤 영향을 미치는가?
  • RQ5같은 모델에 대해 다양한 디바이스에서 정확도의 변동성은 어떤가?

주요 결과

  • AIoTBench는 모바일 AI 추론 벤치마크를 위해 6개 모델을 3개 프레임워크에서 다룬다.
  • 두 개의 통합 지표 VIPS와 VOPS가 정확도 가중치를 가진 엔드투엔드 추론 품질과 처리량을 요약한다.
  • 결과는 장비, 모델, 프레임워크에 따라 성능이 달라지며 모든 조건에서 단일 프레임워크가 지배하지 않는다.
  • 같은 모델과 구현이 다른 디바이스에서 서로 다른 정확도를 보일 수 있다; Oppo R17은 약간의 정확도 편차를 보인다.
  • 다른 디바이스에서 프레임워크 성능 순서가 바뀌는데, 예를 들어 PyTorch Mobile이 일부 디바이스에서 더 빠른 반면 TensorFlow Lite CPU나 NNAPI 위임은 모델에 따라 다르게 작용한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.