How Many Samples are Needed to Learn a Convolutional Neural Network

Open AccessProceedings Article

How Many Samples are Needed to Learn a Convolutional Neural Network

- Vol. 31, pp 371-381

TLDR

The study of rigorously characterizing the sample complexity of estimating CNNs is initiated, showing that for an $m$-dimensional convolutional filter with linear activation acting on a d-dimensional input, the samplecomplexity of achieving population prediction error of $\epsilon$ is $\widetilde{O(m/\ep silon^2)$, whereas the sample-complexity for its FNN counterpart is lower bounded by $\Omega(d/\Epsilon

Abstract:

A widespread folklore for explaining the success of convolutional neural network (CNN) is that CNN is a more compact representation than the fully connected neural network (FNN) and thus requires fewer samples for learning. We initiate the study of rigorously characterizing the sample complexity of learning convolutional neural networks. We show that for learning an m-dimensional convolutional filter with linear activation acting on a d-dimensional input, the sample complexity of achieving population prediction error of ϵ is ˜ O (m/ϵ2) whereas its FNN counterpart needs at least Ω(d/ϵ2) samples. Since m≪d, this result demonstrates the advantage of using CNN. We further consider the sample complexity of learning a one-hidden-layer CNN with linear activation where both the m-dimensional convolutional filter and the r-dimensional output weights are unknown. For this model, we show the sample complexity is ˜ O ((m+r)/ϵ2) when the ratio between the stride size and the filter size is a constant. For both models, we also present lower bounds showing our sample complexities are tight up to logarithmic factors. Our main tools for deriving these results are localized empirical process and a new lemma characterizing the convolutional structure. We believe these tools may inspire further developments in understanding CNN.

How Many Samples are Needed to Learn a Convolutional Neural Network

Citations

Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks

Generalization bounds for deep convolutional neural networks

Using convolutional neural network for predicting cyanobacteria concentrations in river water

Size-free generalization bounds for convolutional neural networks

Why Are Convolutional Nets More Sample-Efficient than Fully-Connected Nets?

References

ImageNet Classification with Deep Convolutional Neural Networks

ImageNet classification with deep convolutional neural networks

Mastering the game of Go with deep neural networks and tree search

Probability Inequalities for sums of Bounded Random Variables

Convolutional networks for images, speech, and time series

Related Papers (5)

An efficient and effective deep convolutional kernel pseudoinverse learner with multi-filter

CryptoDL: Deep Neural Networks over Encrypted Data

ProdSumNet: reducing model parameters in deep neural networks via product-of-sums matrix decompositions

A Fixed-Point Quantization Technique for Convolutional Neural Networks Based on Weight Scaling

Parameter Distribution Balanced CNNs