Harnessing the Power of Infinitely Wide Deep Nets on Small-data Tasks

Open AccessPosted Content

Harnessing the Power of Infinitely Wide Deep Nets on Small-data Tasks

Sanjeev Arora, +5 more

- 03 Oct 2019 -

arXiv: Learning

Chats0

TLDR

Arora et al. as discussed by the authors showed that neural tangent kernels perform strongly on low-data tasks, and compared their performance with the finite-width net derived from the same neural network, showing that NTK's efficacy may trace to lower variance of output.

Abstract:

Recent research shows that the following two models are equivalent: (a) infinitely wide neural networks (NNs) trained under l2 loss by gradient descent with infinitesimally small learning rate (b) kernel regression with respect to so-called Neural Tangent Kernels (NTKs) (Jacot et al., 2018). An efficient algorithm to compute the NTK, as well as its convolutional counterparts, appears in Arora et al. (2019a), which allowed studying performance of infinitely wide nets on datasets like CIFAR-10. However, super-quadratic running time of kernel methods makes them best suited for small-data tasks. We report results suggesting neural tangent kernels perform strongly on low-data tasks. 1. On a standard testbed of classification/regression tasks from the UCI database, NTK SVM beats the previous gold standard, Random Forests (RF), and also the corresponding finite nets. 2. On CIFAR-10 with 10 - 640 training samples, Convolutional NTK consistently beats ResNet-34 by 1% - 3%. 3. On VOC07 testbed for few-shot image classification tasks on ImageNet with transfer learning (Goyal et al., 2019), replacing the linear SVM currently used with a Convolutional NTK SVM consistently improves performance. 4. Comparing the performance of NTK with the finite-width net it was derived from, NTK behavior starts at lower net widths than suggested by theoretical analysis(Arora et al., 2019a). NTK's efficacy may trace to lower variance of output.

Harnessing the Power of Infinitely Wide Deep Nets on Small-data Tasks

Citations

Few-Shot Learning via Learning the Representation, Provably

Neural Tangents: Fast and Easy Infinite Neural Networks in Python

How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks.

What Can Neural Networks Reason About

Differentially Private Learning Needs Better Features (or Much More Data)

References

Deep Residual Learning for Image Recognition

ImageNet: A large-scale hierarchical image database

Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond

Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)

Do we need hundreds of classifiers to solve real world classification problems

Related Papers (5)

Harnessing the Power of Infinitely Wide Deep Nets on Small-data Tasks

Trading representability for scalability: adaptive multi-hyperplane machine for nonlinear classification

Field Support Vector Machines

Support Vector Machines on Large Data Sets: Simple Parallel Approaches

Fast SVM training using approximate extreme points