Paper relaxes SGD privacy and generalization guarantees for non-smooth convex losses.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Optimal private ERM and SCO with subquadratic gradient complexity.
New algorithms for differentially private optimization in convex and non-convex settings with near-optimal rates.
New algorithm for robust high-dimensional linear regression is both fast and statistically optimal.
Given a convex optimization problem and its dual, there are many possible first-order algorithms. In this paper, we show the equivalence between mirror descent algorithms and algorithms generalizing the conditional gradient method. This is done through convex duality, and implies notably that for certain problems, such…
In recent literature, a general two step procedure has been formulated for solving the problem of phase retrieval. First, a spectral technique is used to obtain a constant-error initial estimate, following which, the estimate is refined to arbitrary precision by first-order optimization of a non-convex loss function. N…
In this paper we study the differentially private Empirical Risk Minimization (ERM) problem in different settings. For smooth (strongly) convex loss function with or without (non)-smooth regularization, we give algorithms that achieve either optimal or near optimal utility bounds with less gradient complexity compared …
MARINA-P improves non-smooth federated optimization with adaptive stepsizes.
A new algorithm improves both computational efficiency and statistical optimality for robust low-rank matrix and tensor estimation.
Advances smooth over-parameterization for solving non-smooth optimization problems.
New bounds for online portfolio selection without smoothness assumptions.
In this paper, we develop a novel {\bf ho}moto{\bf p}y {\bf s}moothing (HOPS) algorithm for solving a family of non-smooth problems that is composed of a non-smooth term with an explicit max-structure and a smooth term or a simple non-smooth term whose proximal mapping is easy to compute. The best known iteration compl…
Smoothness analysis of adversarial training reveals constraints cause more non-smoothness.
New algorithms optimize non-smooth, non-convex objectives with improved complexity.
AsylADMM improves gossip-based learning for non-smooth objectives.
We consider the problem of finding local minimizers in non-convex and non-smooth optimization. Under the assumption of strict saddle points, positive results have been derived for first-order methods. We present the first known results for the non-smooth case, which requires different analysis and a different algorithm…
This work speeds up hyperparameter selection for non-smooth convex models using implicit differentiation.
New algorithms minimize dynamic regret for strongly convex losses.
Paper tackles private optimization for non-smooth objectives efficiently.
A new method lifts training of input-convex neural networks to avoid dead weights and plateaued loss.
New methods improve convergence in non-convex non-smooth learning problems.
This paper presents an asynchronous incremental aggregated gradient algorithm and its implementation in a parameter server framework for solving regularized optimization problems. The algorithm can handle both general convex (possibly non-smooth) regularizers and general convex constraints. When the empirical data loss…
Traditional dictionary learning methods are based on quadratic convex loss function and thus are sensitive to outliers. In this paper, we propose a generic framework for robust dictionary learning based on concave losses. We provide results on composition of concave functions, notably regarding super-gradient computati…
New bounds explain deterministic non-smooth deep nets without large Lipschitz constants.
Expanding FCCO to non-smooth weakly-convex problems, improving deep learning performance.
New SPS variant improves non-smooth optimization without small gradients.
New algorithms achieve optimal DP convex optimization with linear time and gradient computations.
The paper explores various stationarity concepts in non-smooth optimization.
We consider the problem of minimizing the sum of an average function of a large number of smooth convex components and a general, possibly non-differentiable, convex function. Although many methods have been proposed to solve this problem with the assumption that the sum is strongly convex, few methods support the non-…
New sampling algorithm for non-smooth potentials.
Predictive models can be used on high-dimensional brain images for diagnosis of a clinical condition. Spatial regularization through structured sparsity offers new perspectives in this context and reduces the risk of overfitting the model while providing interpretable neuroimaging signatures by forcing the solution to …
Classification is the most important process in data analysis. However, due to the inherent non-convex and non-smooth structure of the zero-one loss function of the classification model, various convex surrogate loss functions such as hinge loss, squared hinge loss, logistic loss, and exponential loss are introduced. T…
Wasserstein distributionally robust optimization (DRO) has recently achieved empirical success for various applications in operations research and machine learning, owing partly to its regularization effect. Although connection between Wasserstein DRO and regularization has been established in several settings, existin…
Stochastic Gradient Descent can overfit after just a few passes, contrary to initial expectations.
Boosting is a popular way to derive powerful learners from simpler hypothesis classes. Following previous work (Mason et al., 1999; Friedman, 2000) on general boosting frameworks, we analyze gradient-based descent algorithms for boosting with respect to any convex objective and introduce a new measure of weak learner p…
A new method solves convex optimization on curved spaces.
Optimization algorithms help overparameterized neural networks achieve high performance.
The paper tackles machine unlearning by designing efficient algorithms for adaptive query classes.
A new optimization method, BPM, converges linearly in non-convex, non-smooth problems.
Iterative Hard Thresholding (IHT) is a class of projected gradient descent methods for optimizing sparsity-constrained minimization models, with the best known efficiency and scalability in practice. As far as we know, the existing IHT-style methods are designed for sparse minimization in primal form. It remains open t…
Adjusting the learning rate schedule in stochastic gradient methods is an important unresolved problem which requires tuning in practice. If certain parameters of the loss function such as smoothness or strong convexity constants are known, theoretical learning rate schedules can be applied. However, in practice, such …
The study extends curvature bounds to non-smooth spaces and proves stability of mean curvature.
Paper proves CLT for quantile SGD with constant learning rate.
New algorithm controls linear systems with bandit feedback, achieving optimal regret.
We provide improved convergence rates for various \emph{non-smooth} optimization problems via higher-order accelerated methods. In the case of regression, we achieves an iteration complexity, breaking the barrier so far present for previous methods. We arrive at a similar rate fo…
Many popular statistical models, such as factor and random effects models, give arise a certain type of covariance structures that is a summation of low rank and sparse matrices. This paper introduces a penalized approximation framework to recover such model structures from large covariance matrix estimation. We propos…
We prove that the empirical risk of most well-known loss functions factors into a linear term aggregating all labels with a term that is label free, and can further be expressed by sums of the loss. This holds true even for non-smooth, non-convex losses and in any RKHS. The first term is a (kernel) mean operator --the …
We consider the problem of finding critical points of functions that are non-convex and non-smooth. Studying a fairly broad class of such problems, we analyze the behavior of three gradient-based methods (gradient descent, proximal update, and Frank-Wolfe update). For each of these methods, we establish rates of conver…