SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks

doi:10.1145/3079856.3080254

Proceedings ArticleDOI

SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks

- Vol. 45, Iss: 2, pp 27-40

TLDR

The Sparse CNN (SCNN) accelerator as discussed by the authors employs a dataflow that enables maintaining the sparse weights and activations in a compressed encoding, which eliminates unnecessary data transfers and reduces storage requirements.

Abstract:

Convolutional Neural Networks (CNNs) have emerged as a fundamental technology for machine learning. High performance and extreme energy efficiency are critical for deployments of CNNs, especially in mobile platforms such as autonomous vehicles, cameras, and electronic personal assistants. This paper introduces the Sparse CNN (SCNN) accelerator architecture, which improves performance and energy efficiency by exploiting the zero-valued weights that stem from network pruning during training and zero-valued activations that arise from the common ReLU operator. Specifically, SCNN employs a novel dataflow that enables maintaining the sparse weights and activations in a compressed encoding, which eliminates unnecessary data transfers and reduces storage requirements. Furthermore, the SCNN dataflow facilitates efficient delivery of those weights and activations to a multiplier array, where they are extensively reused; product accumulation is performed in a novel accumulator array. On contemporary neural networks, SCNN can improve both performance and energy by a factor of 2.7x and 2.3x, respectively, over a comparably provisioned dense CNN accelerator.

Citations

PDF

Open Access

More filters

Journal ArticleDOI

Efficient Processing of Deep Neural Networks: A Tutorial and Survey

Vivienne Sze, +3 more

TL;DR: In this paper, the authors provide a comprehensive tutorial and survey about the recent advances toward the goal of enabling efficient processing of DNNs, and discuss various hardware platforms and architectures that support DNN, and highlight key trends in reducing the computation cost of deep neural networks either solely via hardware design changes or via joint hardware and DNN algorithm changes.

...read moreread less

Book ChapterDOI

AMC: AutoML for Model Compression and Acceleration on Mobile Devices

Yihui He, +5 more

TL;DR: This paper proposes AutoML for Model Compression (AMC) which leverages reinforcement learning to efficiently sample the design space and can improve the model compression quality and achieves state-of-the-art model compression results in a fully automated way without any human efforts.

...read moreread less

Posted Content

Efficient Processing of Deep Neural Networks: A Tutorial and Survey

Vivienne Sze, +3 more

- 27 Mar 2017 -

arXiv: Computer Vision and Pattern Recog...

TL;DR: In this article, the authors provide a comprehensive tutorial and survey about the recent advances towards the goal of enabling efficient processing of DNNs, and discuss various hardware platforms and architectures that support deep neural networks.

...read moreread less

Posted Content

To prune, or not to prune: exploring the efficacy of pruning for model compression

Michael H. Zhu, +1 more

- 05 Oct 2017 -

arXiv: Machine Learning

TL;DR: In this article, the authors investigate two distinct paths for model compression within the context of energy-efficient inference in resource-constrained environments and propose a new gradual pruning technique that is simple and straightforward to apply across a variety of models/datasets with minimal tuning.

...read moreread less

Posted Content

AMC: AutoML for Model Compression and Acceleration on Mobile Devices.

Yihui He, +5 more

- 10 Feb 2018 -

arXiv: Computer Vision and Pattern Recog...

TL;DR: This paper proposed AutoML for Model Compression (AMC) which leverages reinforcement learning to provide the model compression policy, which outperforms conventional rule-based compression policy by having higher compression ratio, better preserving the accuracy and freeing human labor.

...read moreread less

Collapse

References

PDF

Open Access

More filters

Journal ArticleDOI

2005 Special Issue: Framewise phoneme classification with bidirectional LSTM and other neural network architectures

Alex Graves, +1 more

- 01 Jun 2005 -

Neural Networks

TL;DR: In this article, a modified, full gradient version of the LSTM learning algorithm was used for framewise phoneme classification, using the TIMIT database, and the results support the view that contextual information is crucial to speech processing, and suggest that bidirectional networks outperform unidirectional ones.

...read moreread less

Posted Content

Deep Speech: Scaling up end-to-end speech recognition

Awni Hannun, +10 more

- 17 Dec 2014 -

arXiv: Computation and Language

TL;DR: Deep Speech, a state-of-the-art speech recognition system developed using end-to-end deep learning, outperforms previously published results on the widely studied Switchboard Hub5'00, achieving 16.0% error on the full test set.

...read moreread less

Proceedings ArticleDOI

DianNao: a small-footprint high-throughput accelerator for ubiquitous machine-learning

Tianshi Chen, +6 more

TL;DR: This study designs an accelerator for large-scale CNNs and DNNs, with a special emphasis on the impact of memory on accelerator design, performance and energy, and shows that it is possible to design an accelerator with a high throughput, capable of performing 452 GOP/s in a small footprint.

...read moreread less

Journal ArticleDOI

Eyeriss: a spatial architecture for energy-efficient dataflow for convolutional neural networks

Yu-Hsin Chen, +2 more

TL;DR: A novel dataflow, called row-stationary (RS), is presented, that minimizes data movement energy consumption on a spatial architecture and can adapt to different CNN shape configurations and reduces all types of data movement through maximally utilizing the processing engine local storage, direct inter-PE communication and spatial parallelism.

...read moreread less

Collapse

SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks

Citations

Efficient Processing of Deep Neural Networks: A Tutorial and Survey

AMC: AutoML for Model Compression and Acceleration on Mobile Devices

Efficient Processing of Deep Neural Networks: A Tutorial and Survey

To prune, or not to prune: exploring the efficacy of pruning for model compression

AMC: AutoML for Model Compression and Acceleration on Mobile Devices.

References

2005 Special Issue: Framewise phoneme classification with bidirectional LSTM and other neural network architectures

End to end speech recognition in English and Mandarin

Deep Speech: Scaling up end-to-end speech recognition

DianNao: a small-footprint high-throughput accelerator for ubiquitous machine-learning

Eyeriss: a spatial architecture for energy-efficient dataflow for convolutional neural networks

Related Papers (5)

Deep Residual Learning for Image Recognition

In-Datacenter Performance Analysis of a Tensor Processing Unit

Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

ImageNet Classification with Deep Convolutional Neural Networks

Very Deep Convolutional Networks for Large-Scale Image Recognition