First order methods can take extremely long to find global minima of non-convex functions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper improves convergence guarantees for SGD algorithms in non-convex smooth functions.
SGD converges to global minimum for structured non-convex functions.
Improved convergence analysis for decentralized non-convex optimization.
Here we study non-convex composite optimization: first, a finite-sum of smooth but non-convex functions, and second, a general function that admits a simple proximal mapping. Most research on stochastic methods for composite optimization assumes convexity or strong convexity of each function. In this paper, we extend t…
Paper estimates differences in multi-attribute Gaussian graphical models using non-convex penalties.
Study on equilibrium with non-convex preferences.
AGGLIO optimizes non-convex functions with local convexity guarantees.
SGD converges to global minimum for certain non-convex functions.
Optimizers find approximate global minima in non-convex problems.
We introduce a novel algorithm for solving learning problems where both the loss function and the regularizer are non-convex but belong to the class of difference of convex (DC) functions. Our contribution is a new general purpose proximal Newton algorithm that is able to deal with such a situation. The algorithm consi…
Unified analysis of multi-attribute graph learning with non-convex penalties.
In this paper, we study a family of non-convex and possibly non-smooth inf-projection minimization problems, where the target objective function is equal to minimization of a joint function over another variable. This problem include difference of convex (DC) functions and a family of bi-convex functions as special cas…
Generalizes smoothness conditions for optimization methods.
We consider the fundamental problem in non-convex optimization of efficiently reaching a stationary point. In contrast to the convex case, in the long history of this basic problem, the only known theoretical results on first-order non-convex optimization remain to be full gradient descent that converges in $O(1/\varep…
Difference of convex (DC) functions cover a broad family of non-convex and possibly non-smooth and non-differentiable functions, and have wide applications in machine learning and statistics. Although deterministic algorithms for DC functions have been extensively studied, stochastic optimization that is more suitable …
SGD's uncertainty quantified in non-convex learning problems.
In this paper, we study stochastic non-convex optimization with non-convex random functions. Recent studies on non-convex optimization revolve around establishing second-order convergence, i.e., converging to a nearly second-order optimal stationary points. However, existing results on stochastic non-convex optimizatio…
Paper shows non-convexity in solutions to Hessian equations.
This paper studies quasar-convex functions to improve optimization methods.
An Euler discretization of the Langevin diffusion is known to converge to the global minimizers of certain convex and non-convex optimization problems. We show that this property holds for any suitably smooth diffusion and that different diffusions are suitable for optimizing different classes of convex and non-convex …
New bounds found for optimizing non-convex functions with noisy data.
Non-convex extremal length found in surface metrics.
AEGD optimizes non-convex functions with dynamic energy updates.
In this paper we develop proximal methods for statistical learning. Proximal point algorithms are useful in statistics and machine learning for obtaining optimization solutions for composite functions. Our approach exploits closed-form solutions of proximal operators and envelope representations based on the Moreau, Fo…
This paper examines the role and efficiency of the non-convex loss functions for binary classification problems. In particular, we investigate how to design a simple and effective boosting algorithm that is robust to the outliers in the data. The analysis of the role of a particular non-convex loss for prediction accur…
Paper tackles non-convex constrained DRO with a stochastic algorithm for large-scale applications.
We analyze stochastic gradient descent for optimizing non-convex functions. In many cases for non-convex functions the goal is to find a reasonable local minimum, and the main concern is that gradient updates are trapped in saddle points. In this paper we identify strict saddle property for non-convex problem that allo…
Heavy Ball method speeds up finding global optima in non-convex problems.
Non-convex optimization is ubiquitous in machine learning. Majorization-Minimization (MM) is a powerful iterative procedure for optimizing non-convex functions that works by optimizing a sequence of bounds on the function. In MM, the bound at each iteration is required to \emph{touch} the objective function at the opti…
This paper explores the non-convex composition optimization in the form including inner and outer finite-sum functions with a large number of component functions. This problem arises in some important applications such as nonlinear embedding and reinforcement learning. Although existing approaches such as stochastic gr…
Several recently proposed architectures of neural networks such as ResNeXt, Inception, Xception, SqueezeNet and Wide ResNet are based on the designing idea of having multiple branches and have demonstrated improved performance in many applications. We show that one cause for such success is due to the fact that the mul…
This paper addresses the problem of sparsity penalized least squares for applications in sparse signal processing, e.g. sparse deconvolution. This paper aims to induce sparsity more strongly than L1 norm regularization, while avoiding non-convex optimization. For this purpose, this paper describes the design and use of…
Develops a new fairness learning approach for multi-task regression models.
SGD and stochastic gradient descent converge at optimal rates for certain non-convex functions.
Optimally shows the distance between perturbed convex functions and their Γ-regularizations.
SGD converges with positive probability for non-convex deep neural networks under specific conditions.
Learning with a {\it convex loss} function has been a dominating paradigm for many years. It remains an interesting question how non-convex loss functions help improve the generalization of learning with broad applicability. In this paper, we study a family of objective functions formed by truncating traditional loss f…
Adaptive momentum method solves non-convex min-max problems.
Paper improves stability analysis of SGD for various loss functions and data distributions.
While optimizing convex objective (loss) functions has been a powerhouse for machine learning for at least two decades, non-convex loss functions have attracted fast growing interests recently, due to many desirable properties such as superior robustness and classification accuracy, compared with their convex counterpa…
In this paper, we consider the convex and non-convex composition problem with the structure , where is the inner function, and is the outer function. We explore the variance reduction based met…
New framework for robust hypothesis testing using Sinkhorn uncertainty sets.
The paper analyzes adaptive algorithms in non-convex optimization landscapes.
Shielded LMC samples from non-convex spaces with repulsive drift.
Improved generalization bounds for SGD in non-convex learning.
Manifolds uniquely identified by boundary distance differences.
We consider the problem of finding local minimizers in non-convex and non-smooth optimization. Under the assumption of strict saddle points, positive results have been derived for first-order methods. We present the first known results for the non-smooth case, which requires different analysis and a different algorithm…