New Max-Plus neural network exploits subgradient sparsity for efficient training.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this paper, we propose a convergent parallel best-response algorithm with the exact line search for the nondifferentiable nonconvex sparsity-regularized rank minimization problem. On the one hand, it exhibits a faster convergence than subgradient algorithms and block coordinate descent algorithms. On the other hand,…
New method estimates sparse mean from noisy data without knowing sparsity level.
New method bounds stochastic subgradient methods with heavy-tailed noise.
The paper guarantees global stability for stochastic subgradient methods in nonsmooth nonconvex optimization.
A distributed subgradient method tackles non-convex optimization problems in networks.
Improved subgradient method tackles ill-conditioned composite optimization problems.
Nesterov's extrapolation improves convergence in nonsmooth optimization.
We study the problem of learning high dimensional regression models regularized by a structured-sparsity-inducing penalty that encodes prior structural information on either input or output sides. We consider two widely adopted types of such penalties as our motivating examples: 1) overlapping group lasso penalty, base…
The paper derives subgradient estimates for a specific nonlinear subparabolic equation on pseudo-Hermitian manifolds.
SNAM improves NAM's accuracy and feature selection via group sparsity.
We show that the Subgradient algorithm is universal for online learning on the simplex in the sense that it simultaneously achieves regret for adversarial costs and pseudo-regret for i.i.d costs. To the best of our knowledge this is the first demonstration of a universal algorithm on the simplex tha…
Proof of convergence for multi-objective optimization using inverse reinforcement learning.
Inexact subgradient methods work well for semialgebraic functions with additive errors.
Unified Lagrangian-based methods for nonsmooth nonconvex optimization.
We study the problem of estimating high-dimensional regression models regularized by a structured sparsity-inducing penalty that encodes prior structural information on either the input or output variables. We consider two widely adopted types of penalties of this kind as motivating examples: (1) the general overlappin…
New algorithms achieve high-probability parameter-free regret in online convex optimization with heavy-tailed data.
Study proves convergence of subgradients for optimal transport-based objectives.
We describe novel subgradient methods for a broad class of matrix optimization problems involving nuclear norm regularization. Unlike existing approaches, our method executes very cheap iterations by combining low-rank stochastic subgradients with efficient incremental SVD updates, made possible by highly optimized and…
Novel coordinate descent (CD) methods are proposed for minimizing nonconvex functions consisting of three terms: (i) a continuously differentiable term, (ii) a simple convex term, and (iii) a concave and continuous term. First, by extending randomized CD to nonsmooth nonconvex settings, we develop a coordinate subgradi…
The paper accelerates ISTA and FISTA algorithms for composite optimization problems.
In this paper we study integer multiplicity rectifiable currents carried by the subgradient (subdifferential) graphs of semi-convex functions on a -dimensional convex domain, and show a weak continuity theorem with respect to pointwise convergence for such currents. As an application, the -Hessian measures are ca…
New algorithms solve large-scale convex regression problems.
Study robust recovery of low-rank matrices from corrupted measurements without rank prior.
We consider the problem of unconstrained online convex optimization (OCO) with sub-exponential noise, a strictly more general problem than the standard OCO. In this setting, the learner receives a subgradient of the loss functions corrupted by sub-exponential noise and strives to achieve optimal regret guarantee, witho…
We develop model-based methods for solving stochastic convex optimization problems, introducing the approximate-proximal point, or aProx, family, which includes stochastic subgradient, proximal point, and bundle methods. When the modeling approaches we propose are appropriately accurate, the methods enjoy stronger conv…
Linear encoding of sparse vectors is widely popular, but is commonly data-independent -- missing any possible extra (but a priori unknown) structure beyond sparsity. In this paper we present a new method to learn linear encoders that adapt to data, while still performing well with the widely used decoder. The …
In this note, we present a new averaging technique for the projected stochastic subgradient method. By using a weighted average with a weight of t+1 for each iterate w_t at iteration t, we obtain the convergence rate of O(1/t) with both an easy proof and an easy implementation. The new scheme is compared empirically to…
Bayesian max-margin models have shown superiority in various practical applications, such as text categorization, collaborative prediction, social network link prediction and crowdsourcing, and they conjoin the flexibility of Bayesian modeling and predictive strengths of max-margin learning. However, Monte Carlo sampli…
Stochastic subgradient descent avoids critical points in definable functions.
This paper proves equivalences of portfolio optimization problems with negative expectile and omega ratio. We derive subgradients for the negative expectile as a function of the portfolio from a known dual representation of expectile and general theory about subgradients of risk measures. We also give an elementary der…
Paper presents an efficient algorithm for learning minimax risk classifiers with large-scale data.
We consider the issue of solution uniqueness for portfolio optimization problem and its inverse for asset returns with a finite number of possible scenarios. The risk is assessed by deviation measures introduced by [Rockafellar et al., Mathematical Programming, Ser. B, 108 (2006), pp. 515-540] instead of variance as in…
New estimator avoids overfitting in convex regression.
In this paper, we define the geometric median of a probability measure on a Riemannian manifold, give its characterization and a natural condition to ensure its uniqueness. In order to calculate the median in practical cases, we also propose a subgradient algorithm and prove its convergence as well as estimating the er…
We relate the minimax game of generative adversarial networks (GANs) to finding the saddle points of the Lagrangian function for a convex optimization problem, where the discriminator outputs and the distribution of generator outputs play the roles of primal variables and dual variables, respectively. This formulation …
In this paper, a new theory is developed for first-order stochastic convex optimization, showing that the global convergence rate is sufficiently quantified by a local growth rate of the objective function in a neighborhood of the optimal solutions. In particular, if the objective function in the -sub…
Given a convex optimization problem and its dual, there are many possible first-order algorithms. In this paper, we show the equivalence between mirror descent algorithms and algorithms generalizing the conditional gradient method. This is done through convex duality, and implies notably that for certain problems, such…
In this paper, we address the problem of reconstructing coverage maps from path-loss measurements in cellular networks. We propose and evaluate two kernel-based adaptive online algorithms as an alternative to typical offline methods. The proposed algorithms are application-tailored extensions of powerful iterative meth…
This work establishes uniform convergence of subdifferentials in stochastic optimization.
New algorithms optimize spectral risk measures, improving interpolation between average and worst-case performance.
We consider the problem of minimizing a convex risk with stochastic subgradients guaranteeing -locally differentially private (-LDP). While it has been shown that stochastic optimization is possible with -LDP via the standard SGD (Song et al., 2013), its convergence rate largely depends on the learning rate, w…
Sparse methods for supervised learning aim at finding good linear predictors from as few variables as possible, i.e., with small cardinality of their supports. This combinatorial selection problem is often turned into a convex optimization problem by replacing the cardinality function by its convex envelope (tightest c…
This paper concerns dictionary learning, i.e., sparse coding, a fundamental representation learning problem. We show that a subgradient descent algorithm, with random initialization, can provably recover orthogonal dictionaries on a natural nonsmooth, nonconvex minimization formulation of the problem, under mi…
We propose a randomized block-coordinate variant of the classic Frank-Wolfe algorithm for convex optimization with block-separable constraints. Despite its lower iteration cost, we show that it achieves a similar convergence rate in duality gap as the full Frank-Wolfe algorithm. We also show that, when applied to the d…
AIHT improves online high-dimensional quantile regression by separating support discovery and refinement.
A new method solves convex optimization on curved spaces.
Paper proposes fast, robust methods for low-rank matrix recovery.