New method bounds stochastic subgradient methods with heavy-tailed noise.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper guarantees global stability for stochastic subgradient methods in nonsmooth nonconvex optimization.
A distributed subgradient method tackles non-convex optimization problems in networks.
Improved subgradient method tackles ill-conditioned composite optimization problems.
New Max-Plus neural network exploits subgradient sparsity for efficient training.
The paper derives subgradient estimates for a specific nonlinear subparabolic equation on pseudo-Hermitian manifolds.
We show that the Subgradient algorithm is universal for online learning on the simplex in the sense that it simultaneously achieves regret for adversarial costs and pseudo-regret for i.i.d costs. To the best of our knowledge this is the first demonstration of a universal algorithm on the simplex tha…
Proof of convergence for multi-objective optimization using inverse reinforcement learning.
Inexact subgradient methods work well for semialgebraic functions with additive errors.
Unified Lagrangian-based methods for nonsmooth nonconvex optimization.
New algorithms achieve high-probability parameter-free regret in online convex optimization with heavy-tailed data.
Study proves convergence of subgradients for optimal transport-based objectives.
We describe novel subgradient methods for a broad class of matrix optimization problems involving nuclear norm regularization. Unlike existing approaches, our method executes very cheap iterations by combining low-rank stochastic subgradients with efficient incremental SVD updates, made possible by highly optimized and…
The paper accelerates ISTA and FISTA algorithms for composite optimization problems.
In this paper we study integer multiplicity rectifiable currents carried by the subgradient (subdifferential) graphs of semi-convex functions on a -dimensional convex domain, and show a weak continuity theorem with respect to pointwise convergence for such currents. As an application, the -Hessian measures are ca…
New algorithms solve large-scale convex regression problems.
Study robust recovery of low-rank matrices from corrupted measurements without rank prior.
We consider the problem of unconstrained online convex optimization (OCO) with sub-exponential noise, a strictly more general problem than the standard OCO. In this setting, the learner receives a subgradient of the loss functions corrupted by sub-exponential noise and strives to achieve optimal regret guarantee, witho…
We develop model-based methods for solving stochastic convex optimization problems, introducing the approximate-proximal point, or aProx, family, which includes stochastic subgradient, proximal point, and bundle methods. When the modeling approaches we propose are appropriately accurate, the methods enjoy stronger conv…
In this note, we present a new averaging technique for the projected stochastic subgradient method. By using a weighted average with a weight of t+1 for each iterate w_t at iteration t, we obtain the convergence rate of O(1/t) with both an easy proof and an easy implementation. The new scheme is compared empirically to…
Bayesian max-margin models have shown superiority in various practical applications, such as text categorization, collaborative prediction, social network link prediction and crowdsourcing, and they conjoin the flexibility of Bayesian modeling and predictive strengths of max-margin learning. However, Monte Carlo sampli…
Stochastic subgradient descent avoids critical points in definable functions.
This paper proves equivalences of portfolio optimization problems with negative expectile and omega ratio. We derive subgradients for the negative expectile as a function of the portfolio from a known dual representation of expectile and general theory about subgradients of risk measures. We also give an elementary der…
Paper presents an efficient algorithm for learning minimax risk classifiers with large-scale data.
We consider the issue of solution uniqueness for portfolio optimization problem and its inverse for asset returns with a finite number of possible scenarios. The risk is assessed by deviation measures introduced by [Rockafellar et al., Mathematical Programming, Ser. B, 108 (2006), pp. 515-540] instead of variance as in…
New estimator avoids overfitting in convex regression.
In this paper, we define the geometric median of a probability measure on a Riemannian manifold, give its characterization and a natural condition to ensure its uniqueness. In order to calculate the median in practical cases, we also propose a subgradient algorithm and prove its convergence as well as estimating the er…
We relate the minimax game of generative adversarial networks (GANs) to finding the saddle points of the Lagrangian function for a convex optimization problem, where the discriminator outputs and the distribution of generator outputs play the roles of primal variables and dual variables, respectively. This formulation …
In this paper, a new theory is developed for first-order stochastic convex optimization, showing that the global convergence rate is sufficiently quantified by a local growth rate of the objective function in a neighborhood of the optimal solutions. In particular, if the objective function in the -sub…
Given a convex optimization problem and its dual, there are many possible first-order algorithms. In this paper, we show the equivalence between mirror descent algorithms and algorithms generalizing the conditional gradient method. This is done through convex duality, and implies notably that for certain problems, such…
This work establishes uniform convergence of subdifferentials in stochastic optimization.
New algorithms optimize spectral risk measures, improving interpolation between average and worst-case performance.
We consider the problem of minimizing a convex risk with stochastic subgradients guaranteeing -locally differentially private (-LDP). While it has been shown that stochastic optimization is possible with -LDP via the standard SGD (Song et al., 2013), its convergence rate largely depends on the learning rate, w…
This paper concerns dictionary learning, i.e., sparse coding, a fundamental representation learning problem. We show that a subgradient descent algorithm, with random initialization, can provably recover orthogonal dictionaries on a natural nonsmooth, nonconvex minimization formulation of the problem, under mi…
In this paper, we propose a convergent parallel best-response algorithm with the exact line search for the nondifferentiable nonconvex sparsity-regularized rank minimization problem. On the one hand, it exhibits a faster convergence than subgradient algorithms and block coordinate descent algorithms. On the other hand,…
We propose a randomized block-coordinate variant of the classic Frank-Wolfe algorithm for convex optimization with block-separable constraints. Despite its lower iteration cost, we show that it achieves a similar convergence rate in duality gap as the full Frank-Wolfe algorithm. We also show that, when applied to the d…
A new method solves convex optimization on curved spaces.
New method estimates sparse mean from noisy data without knowing sparsity level.
Paper proposes fast, robust methods for low-rank matrix recovery.
In this paper we analyze boosting algorithms in linear regression from a new perspective: that of modern first-order methods in convex optimization. We show that classic boosting algorithms in linear regression, namely the incremental forward stagewise algorithm (FS) and least squares boosting (LS-Boost($…
Study on tensor nuclear norm's decomposability and subdifferential.
New method for optimization on Hadamard manifolds with curvature-independent guarantees.
Study on Adam-family methods for nonsmooth optimization with convergence guarantees.
We present new algorithms to compute the mean of a set of empirical probability measures under the optimal transport metric. This mean, known as the Wasserstein barycenter, is the measure that minimizes the sum of its Wasserstein distances to each element in that set. We propose two original algorithms to compute Wasse…
Optimized method tackles convex optimization with heavy-tailed noise.
Nesterov's extrapolation improves convergence in nonsmooth optimization.
We propose graph-dependent implicit regularisation strategies for distributed stochastic subgradient descent (Distributed SGD) for convex problems in multi-agent learning. Under the standard assumptions of convexity, Lipschitz continuity, and smoothness, we establish statistical learning rates that retain, up to logari…
New algorithms accelerate model-based optimization for stochastic problems.