Paper accelerates an optimization algorithm using extrapolation techniques.
problem Optimizing problems with a sum of a differentiable loss and a nonconvex sparsity regularizer.
method Incorporates extrapolation techniques into iteratively reweighted ℓ 1 \ell_1 ℓ 1 algorithms. result Sequence generated clusters at stationary points of the optimization problem.
Proposes a new method for joint sample and feature selection in multi-view data.
problem Cannot detect latent subsets of samples and remove outliers.
method Weighted Sparse Partial Least Squares ( ℓ ∞ / ℓ 0 \ell_\infty/\ell_0 ℓ ∞ / ℓ 0 -wsPLS) method for joint sample and feature selection. result Developed globally convergent algorithm and iterative algorithms for multi-view data fusion.
A new method combines extrapolation and line search for solving nonconvex, nonsmooth optimization problems.
problem Nonconvex, nonsmooth optimization problems in machine learning and image processing.
method Proximal gradient method with extrapolation and line search (PGels).
result The method reduces to existing algorithms under proper parameter choices and converges to stationary points.
PPGD solves nonconvex nonsmooth optimization problems without KL property.
problem Nonconvex and nonsmooth optimization problems in statistics and machine learning.
method Projective Proximal Gradient Descent (PPGD) for solving a class of nonconvex and nonsmooth problems.
result PPGD achieves a fast convergence rate of O(1/k^2) for k ≥ k_0.
Efficient algorithm solves sparse nonconvex regression problems.
problem Sparse nonconvex square-root-loss regression problems.
method Proximal majorization-minimization (PMM) algorithm with sparse semismooth Newton method.
result Converges to a d-stationary point with Kurdyka-Łojasiewicz property.
New algorithm solves nonconvex problems efficiently.
problem Nonconvex and nonsmooth problems in signal processing and machine learning.
method Reweighted Alternating Direction Method of Multipliers with linearization.
result The algorithm globally converges to a critical point.
A new algorithm for training deep neural networks efficiently.
problem Efficient training of deep neural networks due to nonconvex optimization.
method Proximal block coordinate descent (BCD) algorithm based on the Kurdyka-Lojasiewicz (KL) property.
result Global convergence results and competitive efficiency compared to standard optimizers.
Paper shows how KL exponent is preserved via inf-projection for optimization problems.
problem Estimating the KL exponent for optimization problems.
method Shows KL exponent preservation via inf-projection for optimization problems.
result KL exponent is preserved via inf-projection for several important convex optimization models.
Introduces PPMM algorithm for nonconvex robust regression problems.
problem Nonconvex tuning-free robust regression problems.
method PPMM algorithm with inner subproblems solved by SSN-PPA.
result Converges to d-stationary point with KL property.
New convergence analysis for ADAM algorithm in non-convex optimization with adaptive step size.
problem Convergence issues in ADAM algorithm for non-convex optimization.
method Study of ADAM algorithm under bounded adaptive step size assumption, providing safe step sizes.
result Novel first order convergence rate result in deterministic and stochastic contexts.
Study efficient iterative method for distribution matching using sliced optimal transport.
problem Efficiently match distributions using sliced optimal transport.
method Slice-matching scheme based on sliced optimal transport, with quantitative non-asymptotic rates derived.
result Derive quantitative non-asymptotic rates for convergence to target distribution.
The paper develops a convergence framework for inexact nonconvex and nonsmooth algorithms.
problem Tackles convergence of inexact nonconvex and nonsmooth algorithms.
method Promises pseudo sufficient descent and relative error conditions, and assumes continuity and Kurdyka-Lojasiewicz property.
result Proves the convergence of algorithms to critical points under specific conditions.
CR method improves convergence for nonconvex optimization under KL property.
problem Improving convergence rate for nonconvex optimization problems.
method Cubic-regularized Newton's method exploiting Kurdyka-Lojasiewicz (KL) property.
result Asymptotic convergence rates of various optimality measures are fully characterized.
This paper improves inverse problem solving with weakly convex regularisers and proves convergence.
problem Improving solution methods for inverse problems.
method Generalised formulation of convergent regularisation using weakly convex regularisers, and proof of convergence for primal-dual hybrid gradient method.
result Proves convergence of primal-dual hybrid gradient method for variational problems and shows improved performance with IWCNNs.
BCD methods provide provable convergence guarantees for deep learning models.
problem Theoretical convergence guarantees for BCD methods in deep learning.
method Established global convergence rate of O(1/k) for DNN training models.
result Global convergence to a critical point at a rate of O(1/k) for most DNN training models.
New method tackles nonconvex-nonconcave problems with local KL condition.
problem Nonconvex-nonconcave minimax problems under varying KL conditions.
method Inexact proximal gradient method for KL-structured subproblems.
result Complexity guarantees for approximate stationary points.
New analysis reveals batch size effects on stochastic conditional gradient methods.
problem Understanding the role of batch size in stochastic conditional gradient methods.
method Deriving a new analysis focusing on momentum-based stochastic conditional gradient algorithms (e.g., Scion).
result Increasing batch size initially improves optimization accuracy but can degrade performance beyond a critical threshold.
Deep networks converge in direction, with implications for predictions and margins.
problem Understanding convergence and alignment in deep learning networks.
method Developed a theory of unbounded nonsmooth Kurdyka-Łojasiewicz inequalities for functions definable in an o-minimal structure.
result Network weights, predictions, training errors, and margin distribution converge in direction and align with gradient flow.
Paper develops accelerated APCD for nonconvex nonsmooth problems with performance guarantees.
problem Efficient methods for nonconvex nonsmooth optimization problems with performance guarantees.
method Asynchronous Accelerated Proximal Coordinate Descent (AAPCD) for nonsmooth and nonconvex problems.
result AAPCD ensures that every limit point is a critical point and achieves linear and sublinear convergence rates.
Paper proposes a framework and algorithm for model compression in neural networks.
problem Training neural networks with model compression techniques suffers from accuracy loss and convergence issues.
method Holistic framework based on nonconvex optimization, using NN-BCD algorithm with closed-form iteration scheme.
result The proposed algorithm globally converges to a critical point at a rate of O(1/k).
Paper proposes iLPA for solving DC composite optimization problems, with applications to matrix completion with outliers.
problem Solving nonconvex and nonsmooth DC composite optimization problems.
method Inexact linearized proximal algorithm (iLPA) for DC composite optimization problems.
result The iLPA achieves local R-linear convergence rate under the Kurdyka-Łöjasiewicz property.
New algorithm solves ℓ 0 \ell_0 ℓ 0 -norm constrained multilinear logistic regression for tensor data.
problem Non-convex and nonsmooth ℓ 0 \ell_0 ℓ 0 -norm constraints in multilinear logistic regression. method APALM + ^+ + method for globally convergent optimization. result APALM + ^+ + ensures convergence to a first-order critical point. In this paper, we study the Kurdyka-Łojasiewicz (KL) exponent, an important quantity for analyzing the convergence rate of first-order methods. Specifically, we develop various calculus rules to deduce the KL exponent of new (possibly nonconvex and nonsmooth) functions formed from functions with known KL exponents. In …
In this paper, we study the efficiency of a {\bf R}estarted {\bf S}ub{\bf G}radient (RSG) method that periodically restarts the standard subgradient method (SG). We show that, when applied to a broad class of convex optimization problems, RSG method can find an ε ε ε -optimal solution with a lower complexity than the SG m…
The paper proves a margin inequality for separating hyperplanes, useful for analyzing algorithmic bias.
problem Analyzing the implicit bias of algorithms in machine learning.
method Proves a nonsmooth Kurdyka-Lojasiewicz inequality for margin function.
result The bias of algorithm iterates converges at least as fast as the square-root of the margin convergence rate.
Paper analyzes convergence rates of SGD for non-convex functions under various assumptions.
problem Analyzing convergence rates of SGD for non-convex functions.
method Studied convergence properties of Stochastic Gradient Descent (SGD) for invex functions under weaker and stronger hypotheses.
result Derives estimates on the rate of convergence of $J(oldsymbolθ_t)$ to its limit for functions satisfying the Polyak-Lojasiewicz (PL) condition.
In this paper we study nonconvex penalization using Bernstein functions whose first-order derivatives are completely monotone. The Bernstein function can induce a class of nonconvex penalty functions for high-dimensional sparse estimation problems. We derive a thresholding function based on the Bernstein penalty and di…
Paper confirms Thom's conjecture for nonlinear evolutions on manifolds.
problem Thom's gradient conjecture for nonlinear evolution equations.
method Extending and settling the conjecture in infinite dimensional problems using Łojasiewicz, L. Simon, and Kurdyka-Mostowski-Parusinski's foundational works.
result Uniqueness of the limiting direction and characterization of convergence rates for both classical and infinite dimensional settings.
New method solves complex constrained optimization problems.
problem Constrained nonconvex-nonconcave minimax optimization problems.
method Inexact proximal gradient method using sequential convex programming.
result Established complexity guarantees for approximate stationary points.
In this paper, we further study the forward-backward envelope first introduced in [28] and [30] for problems whose objective is the sum of a proper closed convex function and a twice continuously differentiable possibly nonconvex function with Lipschitz continuous gradient. We derive sufficient conditions on the origin…
Bounds on gradient descent and flow paths for convex and nonconvex functions.
problem Understanding the path length of gradient descent and flow curves.
method Analytical derivation of path length bounds for various smooth convex and nonconvex functions.
result Bounds on path length ζ ζ ζ for gradient descent and flow, providing insights into convergence properties. Efficient solver for nonconvex tensor regularization reduces computational cost.
problem Computational inefficiency in extending nonconvex regularization to tensor learning.
method Proximal average algorithm with adaptive momentum, maintaining sparse plus low-rank structure.
result Shows good statistical performance and accuracy on tensor completion problems.
Book introduces deep learning methods with math, theory, and applications.
problem Understanding deep learning algorithms and their mathematical foundations.
method Reviews various ANN architectures and optimization methods, covers theoretical aspects.
result Provides a solid mathematical foundation for deep learning.
New method recovers matrices with nonlinear structures using optimization on Grassmann manifold.
problem Recovering high-rank matrices with nonlinear structures like subspaces or clusters.
method Formulated as rank minimization of a nonlinear feature map, approximated by constrained non-convex optimization on the Grassmann manifold, using Riemannian and alternating minimization schemes.
result Global convergence and worst-case complexity bounds for alternating minimization scheme, leading to unique limit point.
Defines smoothness of definable sets in o-minimal structures.
problem Characterizing smoothness of definable sets in o-minimal structures.
method Characterizes smoothness using tangent cones and metric properties.
result Equivalence of several conditions for C 1 C^1 C 1 smoothness of definable sets. The paper proposes an efficient algorithm for solving Schatten- p p p quasi-norm problems.
problem Finding low-rank solutions of linear inverse problems with Schatten- p p p quasi-norm regularization. method Dynamic proximal gradient algorithm using Cayley transformation and adaptive step size selection.
result The algorithm converges to a stationary point of the objective function under mild assumptions.
SONATA algorithm converges to solutions of nonconvex smooth functions with KL property.
problem Decentralized optimization over networks with nonconvex smooth functions and convex constraints.
method Decentralized gradient-tracking algorithm SONATA under the KL property.
result SONATA converges to stationary solutions at R-linear rate for θ ∈ ( 0 , 1 / 2 ] θ\in (0,1/2] θ ∈ ( 0 , 1/2 ] , sublinear rate for θ ∈ ( 1 / 2 , 1 ) θ\in (1/2,1) θ ∈ ( 1/2 , 1 ) , and R-linear rate for θ = 0 θ=0 θ = 0 . New method improves sampling for weakly log-concave posteriors.
problem Sampling from weakly log-concave posterior distributions.
method Stochastic Langevin Monte Carlo with over-damped diffusion.
result Simulation horizon is ( d log ( n ) 2 ) ( 1 + r ) 2 (d \log(n)^2)^{(1+r)^2} ( d log ( n ) 2 ) ( 1 + r ) 2 with Poisson subsampling. Develops a mean-field theory for multi-head self-attention under cross-entropy training.
problem Mean-field analysis of multi-head self-attention under cross-entropy training.
method Mean-field theory for a simplified single-layer causal multi-head self-attention model.
result Proves a static finite-head approximation bound for the optimal risk.
Refines pDCA_e for DC function minimization, with applications to sparse recovery and outlier detection.
problem Minimizing DC functions with specific properties.
method Refined convergence analysis of pDCA_e algorithm.
result The pDCA_e algorithm converges for level-bounded DC functions without differentiability assumptions.
DS-GDA solves nonconvex-nonconcave problems without regularity conditions.
problem Nonconvex-nonconcave minimax optimization challenges.
method Doubly smoothed gradient descent ascent method (DS-GDA).
result Achieves convergence on various nonconvex-nonconcave problems.
Paper proves SHB convergence with biased gradients and approximate step sizes.
problem Establishing convergence of SHB with biased gradients and approximate step sizes.
method Generalizes SHB convergence conditions for biased gradients, approximate step sizes, and block updating.
result Proves convergence of SHB with new conditions for biased gradients and approximate step sizes.
New method recovers sparse signals from nonlinear observations with robust error bounds.
problem Recovering two sparse vectors from nonlinearly mixed observations with limited data.
method Regularization-based framework combining Huberized data fidelity and generalized folded-concave penalties with a proximal alternating algorithm.
result Estimation error bounds of order σ s log ( n ) / m σ\sqrt{s\log(n)/m} σ s log ( n ) / m at every localized stationary point, with oracle rate σ s / m σ\sqrt{s/m} σ s / m under beta-min condition. New method explores structural sparsity in deep networks efficiently.
problem Learning structural sparsity in over-parameterized deep networks.
method Differential inclusions of inverse scale spaces, coupled with Deep structure splitting Linearized Bregman Iteration (DessiLBI).
result Achieves comparable and better performance in sparse structure exploration than competitive optimizers.
Boosted Difference of Convex Functions Algorithm solves VaR constrained portfolio optimization.
problem Designing VaR optimal portfolios under financial regulations.
method Boosted Difference of Convex Functions Algorithm (BDCA) with a novel line search framework.
result BDCA linearly converges to a Karush-Kuhn-Tucker point for VaR constrained portfolio problems.
New algorithm for private non-convex optimization with optimal rates.
problem Private optimization of non-convex functions under KL condition.
method Variance-reduced gradient descent and proximal point method.
result Achieves nearly optimal rates for excess empirical risk.