A simple proximal stochastic gradient method for nonsmooth nonconvex optimization

Open AccessProceedings Article

A simple proximal stochastic gradient method for nonsmooth nonconvex optimization

- Vol. 31, pp 5569-5579

TLDR

ProxSVRG+ as discussed by the authors is a proximal stochastic gradient algorithm based on variance reduction, which can automatically switch to the faster linear convergence in some regions as long as the objective function satisfies the Polyak-Łojasiewicz condition locally in these regions.

Abstract:

We analyze stochastic gradient algorithms for optimizing nonconvex, nonsmooth finite-sum problems. In particular, the objective function is given by the summation of a differentiable (possibly nonconvex) component, together with a possibly non-differentiable but convex component. We propose a proximal stochastic gradient algorithm based on variance reduction, called ProxSVRG+. Our main contribution lies in the analysis of ProxSVRG+. It recovers several existing convergence results and improves/generalizes them (in terms of the number of stochastic gradient oracle calls and proximal oracle calls). In particular, ProxSVRG+ generalizes the best results given by the SCSG algorithm, recently proposed by [Lei et al., 2017] for the smooth nonconvex case. ProxSVRG+ is also more straightforward than SCSG and yields simpler analysis. Moreover, ProxSVRG+ outperforms the deterministic proximal gradient descent (ProxGD) for a wide range of minibatch sizes, which partially solves an open problem proposed in [Reddi et al., 2016]. Also, ProxSVRG+ uses much less proximal oracle calls than ProxSVRG [Reddi et al., 2016]. Moreover, for nonconvex functions satisfied Polyak-Łojasiewicz condition, we prove that ProxSVRG+ achieves a global linear convergence rate without restart unlike ProxSVRG. Thus, it can automatically switch to the faster linear convergence in some regions as long as the objective function satisfies the PL condition locally in these regions. Finally, we conduct several experiments and the experimental results are consistent with the theoretical results.

A simple proximal stochastic gradient method for nonsmooth nonconvex optimization

Citations

SpiderBoost: A Class of Faster Variance-reduced Algorithms for Nonconvex Optimization.

ProxSARAH: An Efficient Algorithmic Framework for Stochastic Composite Nonconvex Optimization

PAGE: A Simple and Optimal Probabilistic Gradient Estimator for Nonconvex Optimization

An Improved Convergence Analysis of Stochastic Variance-Reduced Policy Gradient

Stochastic AUC Maximization with Deep Neural Networks.

References

Introductory Lectures on Convex Optimization: A Basic Course

Accelerating Stochastic Gradient Descent using Predictive Variance Reduction

SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives

On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima

A proximal stochastic gradient method with progressive variance reduction

Related Papers (5)

Accelerating Stochastic Gradient Descent using Predictive Variance Reduction

SARAH: A Novel Method for Machine Learning Problems Using Stochastic Recursive Gradient

SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives

Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization

Introductory Lectures on Convex Optimization: A Basic Course