First order methods can take extremely long to find global minima of non-convex functions.
problem Finding global minimizers of non-convex functions.
method Designing a family of non-convex functions and using statistical lower bounds for parameter estimation.
result First order methods can take exponential time to converge to a global minimizer.
This paper improves convergence guarantees for SGD algorithms in non-convex smooth functions.
problem Theoretical convergence properties of SGD algorithms for non-convex smooth functions.
method Analysis of SGD algorithms with arbitrary data ordering for non-convex smooth functions.
result Enhanced convergence guarantees for incremental gradient and single shuffle SGD, improving the optimization term of convergence guarantee.
Paper tackles non-convex inf-projection problems with stochastic optimization.
problem Non-convex and possibly non-smooth inf-projection minimization problems.
method Developed stochastic algorithms for finding (nearly) stationary solutions.
result Established first-order convergence for non-convex inf-projection problems.
New diffusions help globally optimize non-convex functions.
problem Optimizing non-convex functions globally.
method Euler discretization of Langevin diffusion.
result Different diffusions optimize different convex and non-convex functions.
SGD method converges to minima for non-convex functions.
problem Non-convex optimization problems in machine learning.
method Stochastic gradient descent method for non-convex objective functions.
result Estimates on the rate of convergence to minima.
New stochastic algorithms solve DC functions and non-convex problems efficiently.
problem Solving non-convex, non-smooth, and non-differentiable functions efficiently.
method Proposed new stochastic optimization algorithms for DC functions and non-convex problems.
result First non-asymptotic convergence for non-convex optimization with general non-convex non-differentiable regularizers.
SGD converges to global minimum for structured non-convex functions.
problem Optimizing non-convex functions using SGD with slow convergence rates.
method Convergence theorems for SGD on structured non-convex functions, including Quasar and PL conditions.
result SGD converges to global minimum for specific non-convex functions under certain conditions.
Improved convergence analysis for decentralized non-convex optimization.
problem Minimizing a sum of smooth non-convex functions over a network.
method Gradient tracking in decentralized stochastic gradient descent (GT-DSGD).
result GT-DSGD achieves network-independent performances matching centralized SGD under certain conditions.
Here we study non-convex composite optimization: first, a finite-sum of smooth but non-convex functions, and second, a general function that admits a simple proximal mapping. Most research on stochastic methods for composite optimization assumes convexity or strong convexity of each function. In this paper, we extend t…
Paper estimates differences in multi-attribute Gaussian graphical models using non-convex penalties.
problem Estimating differences in multi-attribute Gaussian graphical models with similar structure.
method Penalized D-trace loss function with non-convex (log-sum and SCAD) penalties, proximal gradient descent methods.
result Theoretical analysis and numerical examples support consistency in support recovery and estimation.
AGGLIO optimizes non-convex functions with local convexity guarantees.
problem Optimizing non-convex functions with local convexity.
method Stage-wise, graduated optimization technique for locally convex functions.
result Global convergence to the global optimum for non-convex and locally convex objectives.
Study on equilibrium with non-convex preferences.
problem Existence of equilibrium in non-convex preference settings.
method Provided a necessary and sufficient condition for equilibrium existence.
result Standard equilibrium theory cannot be applied to non-convex preferences.
SGD converges to global minimum for certain non-convex functions.
problem Theoretical challenges in optimizing non-convex functions in machine learning.
method Perturbed SGD on a broad class of non-convex functions.
result SGD converges to global minimum for certain non-convex functions.
Optimizers find approximate global minima in non-convex problems.
problem Understanding why local methods solve non-convex optimization problems.
method Formalizing the hypothesis that many local minima are approximately global minima.
result Most local minima of practical non-convex objectives are approximately global minima.
We introduce a novel algorithm for solving learning problems where both the loss function and the regularizer are non-convex but belong to the class of difference of convex (DC) functions. Our contribution is a new general purpose proximal Newton algorithm that is able to deal with such a situation. The algorithm consi…
Unified analysis of multi-attribute graph learning with non-convex penalties.
problem Graph inference from multi-attribute data.
method Penalized log-likelihood objective function with ADMM and local linear approximation.
result Local consistency in support recovery and precision matrix estimation for non-convex penalties.
Generalizes smoothness conditions for optimization methods.
problem Optimization under non-uniform smoothness conditions.
method Develops a new analysis technique for bounding gradients.
result Obtains convergence rates for gradient descent and Nesterov's method.
New algorithm finds local minima in non-convex, non-smooth problems.
problem Finding local minimizers in non-convex and non-smooth optimization.
method Perturbed Proximal Descent, tailored for non-smooth cases.
result First known results for non-smooth optimization.
We consider the fundamental problem in non-convex optimization of efficiently reaching a stationary point. In contrast to the convex case, in the long history of this basic problem, the only known theoretical results on first-order non-convex optimization remain to be full gradient descent that converges in $O(1/\varep…
SGD's uncertainty quantified in non-convex learning problems.
problem Uncertainty quantification in non-convex learning problems.
method Asymptotic normality of SGD iterates and bias characterization.
result SGD iterates are asymptotically normally distributed around the expected value of the invariant distribution.
In this paper, we study stochastic non-convex optimization with non-convex random functions. Recent studies on non-convex optimization revolve around establishing second-order convergence, i.e., converging to a nearly second-order optimal stationary points. However, existing results on stochastic non-convex optimizatio…
Flexible ADMM-based algorithm for non-convex problems with convergence guarantees.
problem Solving non-convex problems with smooth and convex components.
method Incorporates ADMM, SCA, distributed, asynchronous, and inexact gradient methods.
result Establishes first-order convergence rate guarantees under mild assumptions.
Study of Langevin processes and their convergence rates for non-convex problems.
problem Convergence of Langevin processes and SGD for non-convex optimization problems.
method Quantitative analysis of convergence rates for discrete Langevin-like processes.
result The convergence of SGD for non-convex problems depends on the potential function and additive noise.
Develops a new interior-point technique for non-convex optimization problems.
problem Optimization of non-convex functions with singularities.
method Generalized self-concordant Hessian-barrier algorithm.
result Global convergence to approximate stationary points with optimal iteration complexity.
This paper studies quasar-convex functions to improve optimization methods.
problem Improving optimization methods for non-convex functions.
method Study of first order methods for quasar-convex functions.
result Proves complexity upper bounds similar to convex functions.
Paper shows non-convexity in solutions to Hessian equations.
problem Non-convexity of solutions to k-Hessian equations in exterior domains. method Examples and new proof for quasiconvexity of harmonic functions.
result Solutions to k-Hessian equations are not quasiconvex in exterior domains. New bounds found for optimizing non-convex functions with noisy data.
problem Limits of first-order stochastic optimization in non-convex settings.
method Divergence decomposition to construct challenging subclasses.
result Sharp lower bounds on noisy gradient queries for various non-convex classes.
This thesis explores how submodularity aids in optimizing non-convex functions and validating algorithms.
problem Understanding which functions can be optimized efficiently in non-convex settings.
method Introducing continuous submodularity and developing algorithms for maximizing these functions.
result Characterization and optimization of continuous submodular functions with strong guarantees.
Non-convex extremal length found in surface metrics.
problem Extremal length functions on surfaces are not always convex.
method Used harmonic maps to R-trees and minimal surfaces in Rn. result Found measured foliations with non-convex extremal length functions.
SGD converges exponentially fast in non-convex over-parametrized learning.
problem Convergence of SGD in non-convex, over-parametrized learning.
method Analysis of SGD with constant step size for non-convex functions satisfying the PL condition.
result Exponential convergence of SGD for non-convex functions satisfying the PL condition.
AEGD optimizes non-convex functions with dynamic energy updates.
problem Optimizing non-convex functions efficiently and robustly.
method Adaptive Gradient Descent (AEGD) with a dynamically updated energy variable.
result AEGD achieves energy-dependent convergence rates for both convex and non-convex objectives.
Differentiable CEM enables end-to-end learning of non-convex optimization problems.
problem Non-convex optimization of continuous, parameterized objective functions.
method Introducing a differentiable variant of the cross-entropy method (CEM).
result Differentiation of CEM output with respect to parameters enables end-to-end learning.
Deep neural networks with multiple branches are less non-convex, improving performance.
problem Improving neural network performance through multi-branch architectures.
method Quantitative measurement of duality gap for neural networks with multi-branches and various activation functions.
result The duality gap of multi-branch neural networks decreases as the number of branches increases, leading to less non-convex optimization problems.
In this paper we develop proximal methods for statistical learning. Proximal point algorithms are useful in statistics and machine learning for obtaining optimization solutions for composite functions. Our approach exploits closed-form solutions of proximal operators and envelope representations based on the Moreau, Fo…
This paper examines the role and efficiency of the non-convex loss functions for binary classification problems. In particular, we investigate how to design a simple and effective boosting algorithm that is robust to the outliers in the data. The analysis of the role of a particular non-convex loss for prediction accur…
Paper tackles non-convex constrained DRO with a stochastic algorithm for large-scale applications.
problem Training robust models against data distribution shifts with non-convex loss functions.
method Developed a stochastic algorithm for non-convex constrained DRO with a complexity independent of dataset size.
result Algorithm finds ε-stationary points with computational complexity of O(ε^(-3k_*-5)) for general Cressie-Read divergence.
We analyze stochastic gradient descent for optimizing non-convex functions. In many cases for non-convex functions the goal is to find a reasonable local minimum, and the main concern is that gradient updates are trapped in saddle points. In this paper we identify strict saddle property for non-convex problem that allo…
Heavy Ball method speeds up finding global optima in non-convex problems.
problem Finding global optima in non-convex optimization problems.
method Heavy Ball momentum in non-convex optimization.
result Heavy Ball helps iterates enter a benign region faster, containing a global optimal point.
Non-convex optimization is ubiquitous in machine learning. Majorization-Minimization (MM) is a powerful iterative procedure for optimizing non-convex functions that works by optimizing a sequence of bounds on the function. In MM, the bound at each iteration is required to \emph{touch} the objective function at the opti…
This paper explores the non-convex composition optimization in the form including inner and outer finite-sum functions with a large number of component functions. This problem arises in some important applications such as nonlinear embedding and reinforcement learning. Although existing approaches such as stochastic gr…
This paper addresses the problem of sparsity penalized least squares for applications in sparse signal processing, e.g. sparse deconvolution. This paper aims to induce sparsity more strongly than L1 norm regularization, while avoiding non-convex optimization. For this purpose, this paper describes the design and use of…
Develops a new fairness learning approach for multi-task regression models.
problem Fairness in multi-task regression models with biased datasets.
method Uses rank-based non-parametric independence test (Mann Whitney U statistic) and reformulates as non-convex optimization problem.
result Outperforms state-of-the-art methods on fairness metrics.
SGD and stochastic gradient descent converge at optimal rates for certain non-convex functions.
problem Optimal convergence rates for non-convex functions under gradient noise.
method Geometric interpretation of the PL-condition to analyze convergence rates.
result Convergence rates of SGD and stochastic gradient descent match those of strongly convex quadratics.
Optimally shows the distance between perturbed convex functions and their Γ-regularizations.
problem Understanding the difference between perturbed convex functions and their Γ-regularizations.
method Analyzing the compactly supported perturbation and the Γ-regularization of a strictly convex function.
result The optimal estimate of the distance between perturbed convex functions and their Γ-regularizations is shown to be o(ε). SGD converges with positive probability for non-convex deep neural networks under specific conditions.
problem Convergence of SGD for non-convex deep neural networks.
method Established local convergence with positive probability under local Łojasiewicz condition and additional structural assumption.
result SGD converges with positive probability for non-convex deep neural networks under specific conditions.
Learning with a {\it convex loss} function has been a dominating paradigm for many years. It remains an interesting question how non-convex loss functions help improve the generalization of learning with broad applicability. In this paper, we study a family of objective functions formed by truncating traditional loss f…
Adam optimizes non-convex functions, converging to critical points.
problem Finding local minima of non-convex functions.
method Continuous-time model of Adam, ODE approximation, decreasing stepsize.
result Adam converges to critical points of non-convex functions.
Adaptive momentum method solves non-convex min-max problems.
problem Non-convex min-max optimization problems in training generative adversarial networks.
method Proposes an adaptive momentum algorithm for non-convex min-max optimization.
result Establishes non-asymptotic convergence rates for the proposed algorithm.