Learning rates in stochastic neural network training are currently determined a priori to training, using expensive manual or automated iterative tuning. This study proposes gradient-only line searches to resolve the learning rate for neural network training algorithms. Stochastic sub-sampling during training decreases…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
GOLS finds activation functions affect training robustness, especially ReLU.
A new method lifts training of input-convex neural networks to avoid dead weights and plateaued loss.
Consider convex optimization problems subject to a large number of constraints. We focus on stochastic problems in which the objective takes the form of expected values and the feasible set is the intersection of a large number of convex sets. We propose a class of algorithms that perform both stochastic gradient desce…
Stochastic approximation algorithms show exponential progress bounds.
Constructs ε-splitting maps for geodesic balls with non-negative Ricci curvature.
New PG methods tackle nonconvex optimization with auto-conditioned stepsizes.
Mini-batch sub-sampling in neural network training is unavoidable, due to growing data demands, memory-limited computational resources such as graphical processing units (GPUs), and the dynamics of on-line learning. In this study we specifically distinguish between static mini-batch sub-sampled loss functions, where mi…
We explore a new approach for training neural networks where all loss functions are replaced by hard constraints. The same approach is very successful in phase retrieval, where signals are reconstructed from magnitude constraints and general characteristics (sparsity, support, etc.). Instead of taking gradient steps, t…
Paper studies PSGD for constrained optimization problems and its statistical properties.
Equivalence of convex optimization, saddle-point problems, and variational inequalities is a well-established concept. The variational inequality (VI) is a static problem which is studied under dynamical settings using a framework called the projected dynamical system, whose stationary points coincide with the static s…
Optimizes reinsurance and investment strategies to minimize ruin probability.
Interpreting gradient methods as fixed-point iterations, we provide a detailed analysis of those methods for minimizing convex objective functions. Due to their conceptual and algorithmic simplicity, gradient methods are widely used in machine learning for massive data sets (big data). In particular, stochastic gradien…
Improved convergence for nonconvex optimization with dependent data.
A new method improves stochastic gradient descent for faster and more efficient estimation.
Stochastic gradient Langevin dynamics (SGLD) is a computationally efficient sampler for Bayesian posterior inference given a large scale dataset. Although SGLD is designed for unbounded random variables, many practical models incorporate variables with boundaries such as non-negative ones or those in a finite interval.…
Two algorithms solve nonconvex minimax problems with linear constraints, achieving complexity guarantees.
The paper proves conditions for compact Kähler manifolds to be projective or rationally connected.
The paper analyzes GTD algorithms with finite-sample bounds.
New algorithm provably converges to second-order stationary points in NMF.
Efficient boosting method for regression with limited feedback.
Applying a well known result for attracting fixed points of biholomorphisms \cite{RR, V}, we observe that one immediately obtains the following result: if is a complete non-compact gradient Kähler-Ricci soliton which is either steady with positive Ricci curvature so that the scalar curvature attains its maxim…
Spectral Clustering is a popular technique to split data points into groups, especially for complex datasets. The algorithms in the Spectral Clustering family typically consist of multiple separate stages (such as similarity matrix construction, low-dimensional embedding, and K-Means clustering as post processing), whi…
New method for zeroth-order stochastic gradient algorithms provides confidence intervals.
New algorithm solves complex optimization problems without needing projections.
We propose a projected semi-stochastic gradient descent method with mini-batch for improving both the theoretical complexity and practical performance of the general stochastic gradient descent method (SGD). We are able to prove linear convergence under weak strong convexity assumption. This requires no strong convexit…
In this work we introduce a conditional accelerated lazy stochastic gradient descent algorithm with optimal number of calls to a stochastic first-order oracle and convergence rate improving over the projection-free, Online Frank-Wolfe based stochastic gradient descent of Hazan an…
The superior performance of ensemble methods with infinite models are well known. Most of these methods are based on optimization problems in infinite-dimensional spaces with some regularization, for instance, boosting methods and convex neural networks use -regularization with the non-negative constraint. However…
This paper focuses on projection-free methods for solving smooth Online Convex Optimization (OCO) problems. Existing projection-free methods either achieve suboptimal regret bounds or have high per-iteration computational costs. To fill this gap, two efficient projection-free online methods called ORGFW and MORGFW are …
We consider gradient estimates to positive solutions of porous medium equations and fast diffusion equations: associated with the Witten Laplacian on Riemannian manifolds. Under the assumption that the -dimensional Bakry-Emery Ricci curvature is bounded from below, we obtain gradient estimates which…
Eigenfunction gradients on curved spaces imply rigid structure.
The paper explores properties of projections and gradient methods in hyperbolic space forms.
SMAVE optimizes SDR by projecting onto a low-dimensional subspace on a Riemannian manifold.
Constructs new steady gradient Ricci solitons for higher dimensions.
Study on quaternionic bisectional curvature for quaternion-Kähler manifolds.
We consider stochastic strongly convex optimization with a complex inequality constraint. This complex inequality constraint may lead to computationally expensive projections in algorithmic iterations of the stochastic gradient descent~(SGD) methods. To reduce the computation costs pertaining to the projections, we pro…
Study on non-negative solutions for stochastic Volterra equations with jumps.
A neural network approach for feature selection using mutual information.
One of the beauties of the projected gradient descent method lies in its rather simple mechanism and yet stable behavior with inexact, stochastic gradients, which has led to its wide-spread use in many machine learning applications. However, once we replace the projection operator with a simpler linear program, as is d…
New bounds for mixing time and privacy in projected Langevin algorithm and noisy SGD.
Stochastic gradient algorithms estimate the gradient based on only one or a few samples and enjoy low computational cost per iteration. They have been widely used in large-scale optimization problems. However, stochastic gradient algorithms are usually slow to converge and achieve sub-linear convergence rates, due to t…
A new PGA algorithm ensures stable, robust, and noise-immune solutions for non-negative inverse problems.
Paper assesses how pandemic data impacts mortality models.
We propose inertial versions of block coordinate descent methods for solving non-convex non-smooth composite optimization problems. Our methods possess three main advantages compared to current state-of-the-art accelerated first-order methods: (1) they allow using two different extrapolation points to evaluate the grad…
We model how Lipschitz continuity changes during neural network training.
To a complex projective structure on a surface, Thurston associates a locally convex pleated surface. We derive bounds on the geometry of both in terms of the norms and of the quadratic differential of given by the Schwarzian derivative of the associated locally univalent map.…
While stochastic variational inference is relatively well known for scaling inference in Bayesian probabilistic models, related methods also offer ways to circumnavigate the approximation of analytically intractable expectations. The key challenge in either setting is controlling the variance of gradient estimates: rec…
We develop a family of reformulations of an arbitrary consistent linear system into a stochastic problem. The reformulations are governed by two user-defined parameters: a positive definite matrix defining a norm, and an arbitrary discrete or continuous distribution over random matrices. Our reformulation has several e…