To prune, or not to prune: exploring the efficacy of pruning for model compression

Open AccessPosted Content

To prune, or not to prune: exploring the efficacy of pruning for model compression

Michael H. Zhu, +1 more

- 05 Oct 2017 -

arXiv: Machine Learning

Chats0

TLDR

In this article, the authors investigate two distinct paths for model compression within the context of energy-efficient inference in resource-constrained environments and propose a new gradual pruning technique that is simple and straightforward to apply across a variety of models/datasets with minimal tuning.

Abstract:

Model pruning seeks to induce sparsity in a deep neural network's various connection matrices, thereby reducing the number of nonzero-valued parameters in the model. Recent reports (Han et al., 2015; Narang et al., 2017) prune deep networks at the cost of only a marginal loss in accuracy and achieve a sizable reduction in model size. This hints at the possibility that the baseline models in these experiments are perhaps severely over-parameterized at the outset and a viable alternative for model compression might be to simply reduce the number of hidden units while maintaining the model's dense connection structure, exposing a similar trade-off in model size and accuracy. We investigate these two distinct paths for model compression within the context of energy-efficient inference in resource-constrained environments and propose a new gradual pruning technique that is simple and straightforward to apply across a variety of models/datasets with minimal tuning and can be seamlessly incorporated within the training process. We compare the accuracy of large, but pruned models (large-sparse) and their smaller, but dense (small-dense) counterparts with identical memory footprint. Across a broad range of neural network architectures (deep CNNs, stacked LSTM, and seq2seq LSTM models), we find large-sparse models to consistently outperform small-dense models and achieve up to 10x reduction in number of non-zero parameters with minimal loss in accuracy.

Citations

PDF

Open Access

More filters

Journal ArticleDOI

Deep Neural Network Compression by In-Parallel Pruning-Quantization

Frederick Tung, +1 more

- 01 Mar 2020 -

IEEE Transactions on Pattern Analysis an...

TL;DR: A deep network compression algorithm that performs weight pruning and quantization jointly, and in parallel with fine-tuning, that improves the state-of-the-art in network compression on AlexNet, VGGNet, GoogLeNet, and ResNet is proposed.

...read moreread less

Posted Content

What Do Compressed Deep Neural Networks Forget

Sara Hooker, +4 more

- 01 Jan 2020 -

arXiv: Learning

TL;DR: This work provides intuition into the role of capacity in deep neural networks and the trade-offs incurred by compression, and finds that models with radically different numbers of weights have comparable top-line performance metrics but diverge considerably in behavior on a narrow subset of the dataset.

...read moreread less

Proceedings Article

One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers

Ari S. Morcos, +3 more

TL;DR: It is found that, within the natural images domain, winning ticket initializations generalized across a variety of datasets, including Fashion MNIST, SVHN, CIFAR-10/100, ImageNet, and Places365, often achieving performance close to that of winning tickets generated on the same dataset.

...read moreread less

Posted Content

Compressing BERT: Studying the Effects of Weight Pruning on Transfer Learning

Mitchell A. Gordon, +2 more

- 19 Feb 2020 -

arXiv: Computation and Language

TL;DR: It is concluded that BERT can be pruned once during pre-training rather than separately for each task without affecting performance, and that fine-tuning BERT on a specific task does not improve its prunability.

...read moreread less

Posted Content

Training independent subnetworks for robust prediction

Marton Havasi, +7 more

- 13 Oct 2020 -

arXiv: Learning

TL;DR: This work shows that, using a multi-input multi-output (MIMO) configuration, one can utilize a single model's capacity to train multiple subnetworks that independently learn the task at hand, and improves model robustness without increasing compute.

...read moreread less

Collapse

References

PDF

Open Access

More filters

Posted Content

Rethinking the Inception Architecture for Computer Vision

Christian Szegedy, +4 more

- 02 Dec 2015 -

arXiv: Computer Vision and Pattern Recog...

TL;DR: This work is exploring ways to scale up networks in ways that aim at utilizing the added computation as efficiently as possible by suitably factorized convolutions and aggressive regularization.

...read moreread less

Posted Content

MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications

Andrew Howard, +7 more

- 17 Apr 2017 -

arXiv: Computer Vision and Pattern Recog...

TL;DR: This work introduces two simple global hyper-parameters that efficiently trade off between latency and accuracy and demonstrates the effectiveness of MobileNets across a wide range of applications and use cases including object detection, finegrain classification, face attributes and large scale geo-localization.

...read moreread less

Proceedings Article

Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Song Han, +3 more

TL;DR: Deep Compression as mentioned in this paper proposes a three-stage pipeline: pruning, quantization, and Huffman coding to reduce the storage requirement of neural networks by 35x to 49x without affecting their accuracy.

...read moreread less

Posted Content

Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Yonghui Wu, +30 more

- 26 Sep 2016 -

arXiv: Computation and Language

TL;DR: GNMT, Google's Neural Machine Translation system, is presented, which attempts to address many of the weaknesses of conventional phrase-based translation systems and provides a good balance between the flexibility of "character"-delimited models and the efficiency of "word"-delicited models.

...read moreread less

Proceedings Article

Learning both weights and connections for efficient neural networks

Song Han, +3 more

TL;DR: In this paper, the authors proposed a method to reduce the storage and computation required by neural networks by an order of magnitude without affecting their accuracy by learning only the important connections using a three-step method.

...read moreread less

Collapse

To prune, or not to prune: exploring the efficacy of pruning for model compression

Citations

Deep Neural Network Compression by In-Parallel Pruning-Quantization

What Do Compressed Deep Neural Networks Forget

One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers

Compressing BERT: Studying the Effects of Weight Pruning on Transfer Learning

Training independent subnetworks for robust prediction

References

Rethinking the Inception Architecture for Computer Vision

MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications

Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Learning both weights and connections for efficient neural networks

Related Papers (5)

Learning both weights and connections for efficient neural networks

Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Deep Residual Learning for Image Recognition

Optimal Brain Damage

Learning Multiple Layers of Features from Tiny Images

Trending Questions (1)