Convolutional Two-Stream Network Fusion for Video Action Recognition

Open AccessPosted Content

Convolutional Two-Stream Network Fusion for Video Action Recognition

- 22 Apr 2016 -

arXiv: Computer Vision and Pattern Recog...

TLDR

In this paper, a spatial and temporal network can be fused at the last convolution layer without loss of performance, but with a substantial saving in parameters, and furthermore, pooling of abstract convolutional features over spatiotemporal neighbourhoods further boosts performance.

Abstract:

Recent applications of Convolutional Neural Networks (ConvNets) for human action recognition in videos have proposed different solutions for incorporating the appearance and motion information. We study a number of ways of fusing ConvNet towers both spatially and temporally in order to best take advantage of this spatio-temporal information. We make the following findings: (i) that rather than fusing at the softmax layer, a spatial and temporal network can be fused at a convolution layer without loss of performance, but with a substantial saving in parameters; (ii) that it is better to fuse such networks spatially at the last convolutional layer than earlier, and that additionally fusing at the class prediction layer can boost accuracy; finally (iii) that pooling of abstract convolutional features over spatiotemporal neighbourhoods further boosts performance. Based on these studies we propose a new ConvNet architecture for spatiotemporal fusion of video snippets, and evaluate its performance on standard benchmarks where this architecture achieves state-of-the-art results.

Citations

PDF

Open Access

More filters

Journal ArticleDOI

Intelligent Dynamic Gesture Recognition Using CNN Empowered by Edit Distance

Shazia Saqib, +4 more

- 01 Feb 2021 -

Cmc-computers Materials & Continua

Posted Content

Learning to score the figure skating sports videos

Chengming Xu, +5 more

- 08 Feb 2018 -

arXiv: Multimedia

TL;DR: This paper proposes a deep architecture that includes two complementary components, i.e., Self-Attentive LSTM and Multi-scale Convolutional Skip L STM that can efficiently learn the local and global sequential information in each video.

...read moreread less

Journal ArticleDOI

Computer Vision-Based Unobtrusive Physical Activity Monitoring in School by Room-Level Physical Activity Estimation: A Method Proposition

Hans Hõrak

- 28 Aug 2019 -

Information-an International Interdiscip...

TL;DR: The proposed method, if applied as a distributed privacy-preserving sensor system, is argued to be useful for monitoring the spatio-temporal distribution of PA in schools over long periods and assessing the efficiency of school-based PA interventions.

...read moreread less

Proceedings ArticleDOI

Temporal Distance Matrices for Squat Classification

Ryoji Ogata, +3 more

TL;DR: An action dataset of squats where different types of poor form have been annotated with a diversity of users and backgrounds is presented, and a model, based on temporal distance matrices, is proposed, which significantly outperforms existing approaches for the classification task.

...read moreread less

Journal ArticleDOI

FSSCaps-DetCountNet: fuzzy soft sets and CapsNet-based detection and counting network for monitoring animals from aerial images

Divya Meena Sundaram, +1 more

- 19 Jun 2020 -

Journal of Applied Remote Sensing

TL;DR: The experimental results and comparative study with other state-of-the-art conventional models demonstrate the effectiveness and robustness of FSSCaps-DetCountNet for real-time animal detection and counting from aerial images.

...read moreread less

Collapse

References

PDF

Open Access

More filters

Proceedings Article

ImageNet Classification with Deep Convolutional Neural Networks

Alex Krizhevsky, +2 more

TL;DR: The state-of-the-art performance of CNNs was achieved by Deep Convolutional Neural Networks (DCNNs) as discussed by the authors, which consists of five convolutional layers, some of which are followed by max-pooling layers, and three fully-connected layers with a final 1000-way softmax.

...read moreread less

Proceedings Article

Very Deep Convolutional Networks for Large-Scale Image Recognition

Karen Simonyan, +1 more

TL;DR: This work investigates the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting using an architecture with very small convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers.

...read moreread less

Proceedings Article

Very Deep Convolutional Networks for Large-Scale Image Recognition

Karen Simonyan, +1 more

TL;DR: In this paper, the authors investigated the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting and showed that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 layers.

...read moreread less

Proceedings ArticleDOI

Going deeper with convolutions

Christian Szegedy, +8 more

TL;DR: Inception as mentioned in this paper is a deep convolutional neural network architecture that achieves the new state of the art for classification and detection in the ImageNet Large-Scale Visual Recognition Challenge 2014 (ILSVRC14).

...read moreread less

Proceedings Article

Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Sergey Ioffe, +1 more

TL;DR: Applied to a state-of-the-art image classification model, Batch Normalization achieves the same accuracy with 14 times fewer training steps, and beats the original model by a significant margin.

...read moreread less

Collapse

Convolutional Two-Stream Network Fusion for Video Action Recognition

Citations

Intelligent Dynamic Gesture Recognition Using CNN Empowered by Edit Distance

Learning to score the figure skating sports videos

Computer Vision-Based Unobtrusive Physical Activity Monitoring in School by Room-Level Physical Activity Estimation: A Method Proposition

Temporal Distance Matrices for Squat Classification

FSSCaps-DetCountNet: fuzzy soft sets and CapsNet-based detection and counting network for monitoring animals from aerial images

References

ImageNet Classification with Deep Convolutional Neural Networks

Very Deep Convolutional Networks for Large-Scale Image Recognition

Very Deep Convolutional Networks for Large-Scale Image Recognition

Going deeper with convolutions

Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Related Papers (5)

Learning Spatiotemporal Features with 3D Convolutional Networks

UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

Deep Residual Learning for Image Recognition

Large-scale Video Classiﬁcation with Convolutional Neural Networks