New insights into continual learning for deep models, showing convergence issues but local linear solutions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Two algorithms solve nonconvex minimax problems with linear constraints, achieving complexity guarantees.
PPGD solves nonconvex nonsmooth optimization problems without KL property.
Study efficient algorithms for nonconvex optimization with state-dependent Markov data.
Paper solves robust multi-dimensional scaling with accelerated projections.
New algorithms solve complex minimax problems without needing derivatives.
New PG methods tackle nonconvex optimization with auto-conditioned stepsizes.
OLLA framework efficiently samples from constrained distributions with nonconvex constraints.
Unified framework for constrained diffusion models on nonconvex sets with efficient landing mechanism.
Nonnegative low-rank matrix recovery can have spurious local minima.
Paper analyzes robust matrix completion with efficient nonconvex method and leave-one-out analysis.
Recent years have seen a flurry of activities in designing provably efficient nonconvex procedures for solving statistical estimation problems. Due to the highly nonconvex nature of the empirical loss, state-of-the-art procedures often require proper regularization (e.g. trimming, regularized cost, projection) in order…
We study Frank-Wolfe methods for nonconvex stochastic and finite-sum optimization problems. Frank-Wolfe methods (in the convex case) have gained tremendous recent interest in machine learning and optimization communities due to their projection-free property and their ability to exploit structured constraints. However,…
Improved convergence for nonconvex optimization with dependent data.
In this paper, the estimation problem for sparse reduced rank regression (SRRR) model is considered. The SRRR model is widely used for dimension reduction and variable selection with applications in signal processing, econometrics, etc. The problem is formulated to minimize the least squares loss with a sparsity-induci…
Python package for projecting onto quadratic hypersurfaces.
Gradient descent with noise converges to a unique optimum in nonconvex matrix factorization.
Schedule-free SGD is optimal for nonconvex optimization problems.
We consider compressed sensing formulated as a minimization problem of nonconvex sparse penalties, Smoothly Clipped Absolute deviation (SCAD) and Minimax Concave Penalty (MCP). The nonconvexity of these penalties is controlled by nonconvexity parameters, and L1 penalty is contained as a limit with respect to these para…
Paper proposes robust tensor regression method for tensor data analysis.
We propose a generic framework based on a new stochastic variance-reduced gradient descent algorithm for accelerating nonconvex low-rank matrix recovery. Starting from an appropriate initial estimator, our proposed algorithm performs projected gradient descent based on a novel semi-stochastic gradient specifically desi…
BMM algorithm improves convergence for nonconvex optimization problems.
This paper analyzes OGDA and EG methods for nonconvex minimax problems.
Studied SGD convergence under weak conditions.
Paper develops methods for non-quadratic loss low-rank matrix recovery.
We present a unified framework for low-rank matrix estimation with nonconvex penalties. We first prove that the proposed estimator attains a faster statistical rate than the traditional low-rank matrix estimator with nuclear norm penalty. Moreover, we rigorously show that under a certain condition on the magnitude of t…
As surrogate functions of -norm, many nonconvex penalty functions have been proposed to enhance the sparse vector recovery. It is easy to extend these nonconvex penalty functions on singular values of a matrix to enhance low-rank matrix recovery. However, different from convex optimization, solving the nonconvex l…
Study uncovers statistical optimality of nonconvex tensor completion methods.
New method solves subspace optimization problems efficiently.
Although the standard formulations of prediction problems involve fully-observed and noiseless data drawn in an i.i.d. manner, many applications involve noisy and/or missing data, possibly involving dependence, as well. We study these issues in the context of high-dimensional sparse linear regression, and propose novel…
A new model for dynamic covariance recovery in neuroimaging data.
The problem of finding the sparsest vector (direction) in a low dimensional subspace can be considered as a homogeneous variant of the sparse recovery problem, which finds applications in robust subspace recovery, dictionary learning, sparse blind deconvolution, and many other problems in signal processing and machine …
New algorithm for nonconvex optimization on constrained Riemannian manifolds converges quickly.
We provide novel theoretical results regarding local optima of regularized -estimators, allowing for nonconvexity in both loss and penalty functions. Under restricted strong convexity on the loss and suitable regularity conditions on the penalty, we prove that \emph{any stationary point} of the composite objective f…
Radial Basis Functions Neural Networks (RBFNNs) are tools widely used in regression problems. One of their principal drawbacks is that the formulation corresponding to the training with the supervision of both the centers and the weights is a highly non-convex optimization problem, which leads to some fundamentally dif…
Large learning rates lead to various implicit biases in nonconvex optimization.
In this paper we study nonconvex penalization using Bernstein functions whose first-order derivatives are completely monotone. The Bernstein function can induce a class of nonconvex penalty functions for high-dimensional sparse estimation problems. We derive a thresholding function based on the Bernstein penalty and di…
The paper analyzes PPM for nonconvex-nonconcave problems, identifying three regions with varying convergence guarantees.
MAML optimizes shared priors for subtasks in a nonconvex meta-objective.
Multiview representation learning is very popular for latent factor analysis. It naturally arises in many data analysis, machine learning, and information retrieval applications to model dependent structures among multiple data sources. For computational convenience, existing approaches usually formulate the multiview …
New SGD analysis for nonconvex optimization finds optimal rates.
New methods improve online matrix optimization with reduced computational cost.
We propose a nonconvex estimator for joint multivariate regression and precision matrix estimation in the high dimensional regime, under sparsity constraints. A gradient descent algorithm with hard thresholding is developed to solve the nonconvex estimator, and it attains a linear rate of convergence to the true regres…
New method reduces communication costs in distributed nonconvex optimization.
We study finite-sum nonconvex optimization problems, where the objective function is an average of nonconvex functions. We propose a new stochastic gradient descent algorithm based on nested variance reduction. Compared with conventional stochastic variance reduced gradient (SVRG) algorithm that uses two reference …
We adapt the Douglas-Rachford (DR) splitting method to solve nonconvex feasibility problems by studying this method for a class of nonconvex optimization problem. While the convergence properties of the method for convex problems have been well studied, far less is known in the nonconvex setting. In this paper, for the…
Momentum Stochastic Gradient Descent (MSGD) algorithm has been widely applied to many nonconvex optimization problems in machine learning, e.g., training deep neural networks, variational Bayesian inference, and etc. Despite its empirical success, there is still a lack of theoretical understanding of convergence proper…
Replica exchange Langevin diffusion accelerates nonconvex optimization.