Two classes of methods have been proposed for escaping from saddle points with one using the second-order information carried by the Hessian and the other adding the noise into the first-order information. The existing analysis for algorithms using noise in the first-order information is quite involved and hides the es…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithm reduces online decision-making regret with efficient LP re-solving and parallel first-order method.
Efficient algorithm for contextual bandits with first-order guarantees.
The filtering-clustering models, including trend filtering and convex clustering, have become an important source of ideas and modeling tools in machine learning and related fields. The statistical guarantee of optimal solutions in these models has been extensively studied yet the investigations on the computational as…
CEFOL uses deep learning for dynamic programming with recursive utility.
Efficient RNN algorithm guarantees convergence in online learning.
Improved first-order algorithm for entropy regularized OT with faster convergence.
New algorithms optimize constrained problems faster, avoiding full set optimization.
New memory-query tradeoffs for convex optimization algorithms.
New inequalities help optimize first-order algorithms for statistical risk analysis.
Memory-constrained algorithms need superlinear memory for efficient convex optimization.
Unified bounds for iterative algorithms with Gaussian data matrices.
This paper improves online learning algorithms for LP problems, achieving better regret bounds.
Paper develops fast method for computing optimal transport.
Generalizes Hamiltonian theory for variational problems, applied to first order gravity.
New algorithm for safer machine learning with different testing and training distributions.
BMM algorithm improves convergence for nonconvex optimization problems.
We establish that first-order methods avoid saddle points for almost all initializations. Our results apply to a wide variety of first-order methods, including gradient descent, block coordinate descent, mirror descent and variants thereof. The connecting thread is that such algorithms can be studied from a dynamical s…
We propose a reduction for non-convex optimization that can (1) turn an stationary-point finding algorithm into an local-minimum finding one, and (2) replace the Hessian-vector product computations with only gradient computations. It works both in the stochastic and the deterministic settings, without hurting the algor…
Geodesic convexity generalizes the notion of (vector space) convexity to nonlinear metric spaces. But unlike convex optimization, geodesically convex (g-convex) optimization is much less developed. In this paper we contribute to the understanding of g-convex optimization by developing iteration complexity analysis for …
We analyze stochastic gradient algorithms for optimizing nonconvex problems. In particular, our goal is to find local minima (second-order stationary points) instead of just finding first-order stationary points which may be some bad unstable saddle points. We show that a simple perturbed version of stochastic recursiv…
A new algorithm solves bilevel optimization with linear constraints.
We study distributed optimization algorithms for minimizing the average of convex functions. The applications include empirical risk minimization problems in statistical machine learning where the datasets are large and have to be stored on different machines. We design a distributed stochastic variance reduced gradien…
In this paper, we consider first-order convergence theory and algorithms for solving a class of non-convex non-concave min-max saddle-point problems, whose objective function is weakly convex in the variables of minimization and weakly concave in the variables of maximization. It has many important applications in mach…
In this paper, we study optimization methods consisting of iteratively minimizing surrogates of an objective function. By proposing several algorithmic variants and simple convergence analyses, we make two main contributions. First, we provide a unified viewpoint for several first-order optimization techniques such as …
We present a predictor-corrector framework, called PicCoLO, that can transform a first-order model-free reinforcement or imitation learning algorithm into a new hybrid method that leverages predictive models to accelerate policy learning. The new "PicCoLOed" algorithm optimizes a policy by recursively repeating two ste…
New method for tuning Graphical Lasso hyperparameters.
Recent applications that arise in machine learning have surged significant interest in solving min-max saddle point games. This problem has been extensively studied in the convex-concave regime for which a global equilibrium solution can be computed efficiently. In this paper, we study the problem in the non-convex reg…
New methods boost first-order optimization with faster rates.
Paper proposes an algorithm to solve complex minimax problems efficiently.
New framework tackles bi-level optimization without LLS condition.
First order methods can take extremely long to find global minima of non-convex functions.
Novel methods for accelerating optimization in complex bilevel and minimax problems.
New lower bounds for bilevel optimization with first-order oracles.
Novel BSG method for efficient stochastic optimization.
LMC algorithm converges to target in Chi-squared and Renyi divergence.
New methods solve optimization problems with heavy-tailed noise, improving upon existing complexity bounds.
In this paper, we propose a new technique named \textit{Stochastic Path-Integrated Differential EstimatoR} (SPIDER), which can be used to track many deterministic quantities of interest with significantly reduced computational cost. We apply SPIDER to two tasks, namely the stochastic first-order and zeroth-order method…
New methods bound estimation error in high-dimensional statistical problems.
A new method speeds up quantum state estimation.
We consider empirical risk minimization of linear predictors with convex loss functions. Such problems can be reformulated as convex-concave saddle point problems, and thus are well suitable for primal-dual first-order algorithms. However, primal-dual algorithms often require explicit strongly convex regularization in …
Adaptive learning rate algorithms such as RMSProp are widely used for training deep neural networks. RMSProp offers efficient training since it uses first order gradients to approximate Hessian-based preconditioning. However, since the first order gradients include noise caused by stochastic optimization, the approxima…
Information geometry applies concepts in differential geometry to probability and statistics and is especially useful for parameter estimation in exponential families where parameters are known to lie on a Riemannian manifold. Connections between the geometric properties of the induced manifold and statistical properti…
A new first-order sampler improves diffusion probabilistic model sampling quality.
New method predicts state evolution for non-first-order algorithms on nonconvex problems.
Quadratic memory is essential for optimal convex optimization queries.
Standard gradient descent methods are susceptible to a range of issues that can impede training, such as high correlations and different scaling in parameter space.These difficulties can be addressed by second-order approaches that apply a pre-conditioning matrix to the gradient to improve convergence. Unfortunately, s…
We provide two fundamental results on the population (infinite-sample) likelihood function of Gaussian mixture models with components. Our first main result shows that the population likelihood function has bad local maxima even in the special case of equally-weighted mixtures of well-separated and spherical…