Skip to main content
QUICK REVIEW

[Paper Review] ActivityNet Challenge 2017 Summary

Bernard Ghanem, Juan Carlos Niebles|arXiv (Cornell University)|Oct 22, 2017
Human Pose and Action RecognitionComputer Science2 references50 citations
TL;DR

A summary of the 2017 ActivityNet Challenge results across tasks, including top submissions and their performance metrics, with notes on methodologies such as feature fusion, two-stream networks, and temporal proposals.

ABSTRACT

The ActivityNet Large Scale Activity Recognition Challenge 2017 Summary: results and challenge participants papers.

Motivation & Objective

  • Stimulate development of improved human activity understanding algorithms for large-scale, untrimmed videos.
  • Present top-performing submissions and their methodologies across all ActivityNet Challenge tasks.
  • Highlight the role of multi-modal features and fusion strategies in advancing performance.

Proposed method

  • Report top-3 submissions for each task and summarize innovative approaches.
  • Present fusion strategies (e.g., CNN, MBH, C3D; weight/hard voting) and temporal models (two-stream, LSTM, TSN).
  • Describe specific model architectures and pipelines used by leading teams (e.g., untrimmed video classification fusion, temporal action proposals with 3D CNNs).
  • Include details on data augmentation, pretraining, and re-ranking strategies when provided.
  • Provide a consolidated view of performance metrics to compare approaches across tasks.

Experimental results

Research questions

  • RQ1What are the leading methods and architectures that performed best on untrimmed video classification and related ActivityNet tasks in 2017?
  • RQ2How do feature fusion and temporal modeling affect performance across untrimmed and trimmed video action recognition?
  • RQ3What are the top-performing approaches for temporal action proposals and dense captions within ActivityNet 2017?
  • RQ4How do data augmentation and class-wise re refinement impact results on challenging, real-world video data?

Key findings

  • Top-3 results for Task 1 (Untrimmed Video Classification): 8.8% top-1 error (IBUG); 9.8% (CHUK, ETHZ, SIAT); 18.9% (Oxford Brookes University & Disney Research).
  • Top-3 results for Task 2 (Trimmed Action Recognition): 12.4% average error (Tsinghua & Baidu); 13.9% (CHUK, ETHZ, SIAT); 14.4% (TwentyBN).
  • Top-3 results for Task 3 (Temporal Action Proposals): AUC of 64.80 (SJTU & Columbia); 64.18% (MSRA); 61.56% (UMD).
  • Top-3 results for Task 4 (Temporal Action Localization): Average mAP of 33.40% (SJTU & Columbia); 31.86% (CHUK, ETHZ, SIAT); 31.82% (IC).
  • Top-3 results for Task 5 (Dense-Captioning Events in Videos): Average Meteor of 12.84 (MSRA); 9.87% (University of Science and Technology of China); 9.61% (RUC & CMU).
  • Several submissions demonstrated that combining multiple feature streams (e.g., CNN, MBH, C3D) and employing fusion strategies (weight and hard voting) can significantly improve untrimmed video classification performance.
  • Innovative approaches highlighted include human/ object attention, class-wise re refinement, two-stream architectures, and multi-scale attention mechanisms.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.