New method finds optimal learning rates for neural nets.
problem Finding optimal learning rates in stochastic neural networks.
method Gradient-only line searches using Non-negative Associative Gradient Projection Points (NN-GPPs).
result Learning rates can be reliably resolved as step sizes along search directions.
GOLS finds activation functions affect training robustness, especially ReLU.
problem Investigate how different activation functions impact GOLS in neural network training.
method Identify SNN-GPPs for GOLS, analyze activation function effects on gradient continuity.
result GOLS robust for most activation functions but sensitive to ReLU.
This study addresses the challenges of dynamic mini-batch sub-sampling in neural network training.
problem Challenges in training neural networks due to dynamic mini-batch sub-sampling.
method Distinguishes between static and dynamic sub-sampling, recasting optimization to find SNN-GPPs.
result SNN-GPPs are less susceptible to sub-sampling-induced discontinuities and better approximate true optima.
A new method lifts training of input-convex neural networks to avoid dead weights and plateaued loss.
problem Training input-convex neural networks with non-negative weights.
method Introduces a hypernetwork that emits non-negative weights from a summary of the input batch, adding stochasticity to soften the loss landscape.
result The lift method achieves lower test loss than projected gradient descent and direct softplus reparametrization.
Consider convex optimization problems subject to a large number of constraints. We focus on stochastic problems in which the objective takes the form of expected values and the feasible set is the intersection of a large number of convex sets. We propose a class of algorithms that perform both stochastic gradient desce…
Constructs ε-splitting maps for geodesic balls with non-negative Ricci curvature.
problem Constructing ε-splitting maps for geodesic balls with non-negative Ricci curvature.
method Induction and stratified almost Gou-Gu Theorem for finding directional points; error estimates for projections.
result Constructs ε-splitting maps on concentric geodesic balls with uniformly small radius. New PG methods tackle nonconvex optimization with auto-conditioned stepsizes.
problem Optimizing nonconvex functions over convex sets.
method Auto-conditioned projected gradient (AC-PG) methods and stochastic variants.
result Achieved optimal iteration complexity for finding approximate stationary points.
This paper addresses sampling from bounded distributions using SGLD.
problem Sampling from models with bounded variables using SGLD.
method Introduces and evaluates various mapping techniques to transform unbounded samples into bounded ones.
result Invertible Lipschitz mappings overcame the pitfalls of existing methods and achieved weak convergence.
Paper studies PSGD for constrained optimization problems and its statistical properties.
problem Online inference for constrained optimization problems.
method Stochastic gradient descent with projection (PSGD) for constrained optimization.
result Limiting distribution of PSGD-based estimates under linear-equality constraints.
New method trains neural nets without loss functions.
problem Training neural networks efficiently and without loss functions.
method Optimizer RRR derives steps from projections to local constraints, not gradients.
result Success in phase retrieval and neural networks, with novel partitioning of projections.
Proposes RNSE for clustering with adaptive similarity matrix learning.
problem Sub-optimal results due to mismatch between stages in Spectral Clustering.
method End-to-end single-stage learning with adaptive similarity matrix and non-negative constraints.
result Superior clustering performance on synthetic and real-world datasets.
Optimizes reinsurance and investment strategies to minimize ruin probability.
problem Optimizing reinsurance and investment strategies to minimize ruin probability.
method Stochastic projected gradient method based on Malliavin calculus.
result Effectiveness of the proposed method demonstrated through numerical experiments.
Improved convergence for nonconvex optimization with dependent data.
problem Constrained smooth nonconvex optimization with dependent data.
method Stochastic projected gradient methods under a general dependent data sampling scheme.
result Achieved worst-case rate of convergence ildeO(t−1/4) and complexity ildeO(ε−4). A new method improves stochastic gradient descent for faster and more efficient estimation.
problem Efficient and fast parametric estimation methods.
method Projected stochastic gradient descent corrected by Fisher scoring.
result The method is faster and more efficient than traditional methods.
Two new methods reduce OCO problem complexity without projections.
problem Efficiently solving smooth Online Convex Optimization problems without projections.
method ORGFW and MORGFW methods using recursive gradient estimation.
result Achieve optimal regret bounds with low computational costs.
Two algorithms solve nonconvex minimax problems with linear constraints, achieving complexity guarantees.
problem Nonconvex minimax problems with coupled linear constraints.
method Zeroth-order primal-dual alternating projected gradient (ZO-PDAPG) and zeroth-order regularized momentum primal-dual projected gradient (ZO-RMPDPG) algorithms.
result Iteration complexity guarantees for solving nonconvex-(strongly) concave minimax problems with coupled linear constraints.
The paper proves conditions for compact Kähler manifolds to be projective or rationally connected.
problem Conditions for compact Kähler manifolds to be projective or rationally connected.
method Proves conditions using quasi-positive and non-negative curvature.
result Compact Kähler manifolds satisfying certain curvature conditions are projective or rationally connected.
Stochastic approximation algorithms show exponential progress bounds.
problem Analyzing the convergence of stochastic approximation algorithms.
method Developed geometric ergodicity proofs to establish exponential concentration bounds.
result Proved faster convergence rates for specific algorithms.
The paper analyzes GTD algorithms with finite-sample bounds.
problem Convergence rate analysis of GTD family of algorithms.
method Formulated as stochastic gradient algorithms and analyzed using saddle-point error.
result Obtained finite-sample bounds on GTD performance.
New algorithm provably converges to second-order stationary points in NMF.
problem Understanding convergence to local minima in NMF.
method Multiplicative weight update dynamics, concurrent updates, and simplex reduction.
result Provable convergence to second-order stationary points.
Efficient boosting method for regression with limited feedback.
problem Online boosting for regression tasks with noisy multi-point bandit feedback.
method Efficient regret minimization method with online boosting algorithm and projection-free online convex optimization.
result Improved state-of-the-art guarantees in efficiency.
Applying a well known result for attracting fixed points of biholomorphisms \cite{RR, V}, we observe that one immediately obtains the following result: if (Mn,g) is a complete non-compact gradient Kähler-Ricci soliton which is either steady with positive Ricci curvature so that the scalar curvature attains its maxim…
Proposes a new algorithm to improve Frank-Wolfe method efficiency.
problem Improves Frank-Wolfe method's stability and efficiency.
method Introduces 1-SFW, a one-sample stochastic Frank-Wolfe algorithm.
result Achieves optimal convergence rate and first-order stationary point.
New algorithm solves complex optimization problems without needing projections.
problem Optimizing nested functions under convex constraints with noisy evaluations.
method Projection-free conditional gradient-type algorithm for smooth stochastic multi-level composition optimization.
result The algorithm achieves ε-stationary solutions with complexity bounds independent of ε and T. We propose a projected semi-stochastic gradient descent method with mini-batch for improving both the theoretical complexity and practical performance of the general stochastic gradient descent method (SGD). We are able to prove linear convergence under weak strong convexity assumption. This requires no strong convexit…
The superior performance of ensemble methods with infinite models are well known. Most of these methods are based on optimization problems in infinite-dimensional spaces with some regularization, for instance, boosting methods and convex neural networks use L1-regularization with the non-negative constraint. However…
In this work we introduce a conditional accelerated lazy stochastic gradient descent algorithm with optimal number of calls to a stochastic first-order oracle and convergence rate O(ε21) improving over the projection-free, Online Frank-Wolfe based stochastic gradient descent of Hazan an…
Eigenfunction gradients on curved spaces imply rigid structure.
problem Eigenfunction gradient estimates on curved manifolds.
method Sharp Li-Yau type gradient estimates for Neumann or Dirichlet eigenfunctions.
result Compact manifolds with specific curvature properties are rigidly structured.
Constructs new steady gradient Ricci solitons for higher dimensions.
problem Finding new steady gradient Ricci solitons with non-negative curvature.
method Constructing continuous families of Ricci flows from spherical polyhedra, proving stability.
result Produces new examples of steady gradient Ricci solitons for n≥4. The paper explores properties of projections and gradient methods in hyperbolic space forms.
problem Optimization problems in hyperbolic space forms.
method Intrinsic κ-projection and gradient projection methods.
result Every accumulation point of the sequence generated by the gradient projection method is a stationary point.
SMAVE optimizes SDR by projecting onto a low-dimensional subspace on a Riemannian manifold.
problem High-dimensional regression challenges due to the curse of dimensionality.
method SMAVE combines nearest-neighbor localization and Riemannian stochastic gradient ascent.
result SMAVE achieves almost-sure convergence and matches RMAVE's synthetic subspace recovery rate.
Study on quaternionic bisectional curvature for quaternion-Kähler manifolds.
problem Characterize quaternionic bisectional curvature on quaternion-Kähler manifolds.
method Analyzing properties of quaternionic bisectional curvature on specific manifolds.
result Non-negative quaternionic bisectional curvature is only on quaternionic projective space.
We consider stochastic strongly convex optimization with a complex inequality constraint. This complex inequality constraint may lead to computationally expensive projections in algorithmic iterations of the stochastic gradient descent~(SGD) methods. To reduce the computation costs pertaining to the projections, we pro…
Study on non-negative solutions for stochastic Volterra equations with jumps.
problem Existence and uniqueness of non-negative solutions for stochastic Volterra equations with jumps and non-Lipschitz coefficients.
method Developed a nonnegative approximation approach and used Yamada--Watanabe approximation technique for convergence proof.
result Established conditions for strong existence and pathwise uniqueness of non-negative solutions.
Stochastic gradient algorithms estimate the gradient based on only one or a few samples and enjoy low computational cost per iteration. They have been widely used in large-scale optimization problems. However, stochastic gradient algorithms are usually slow to converge and achieve sub-linear convergence rates, due to t…
A new PGA algorithm ensures stable, robust, and noise-immune solutions for non-negative inverse problems.
problem Stable convergence and suboptimal solutions in inverse problems due to negative values and high sensitivity to hyperparameters.
method A novel multiplicative update proximal gradient algorithm (SSO-PGA) that enforces non-negativity and boundedness through a learnable sigmoid-based operator.
result Significantly surpasses traditional PGA and other state-of-the-art algorithms in performance and stability.
We model how Lipschitz continuity changes during neural network training.
problem Understanding how Lipschitz continuity evolves during training.
method We use a system of stochastic differential equations to capture the dynamics of Lipschitz continuity under SGD.
result We identify three factors driving the evolution of Lipschitz continuity: gradient flow projection, gradient noise, and Hessian projection.
SSRGD finds local minima in nonconvex problems with simple gradient updates.
problem Finding local minima in nonconvex optimization problems.
method Simple perturbed stochastic recursive gradient descent (SSRGD).
result SSRGD finds (ε,δ)-second-order stationary points efficiently. While stochastic variational inference is relatively well known for scaling inference in Bayesian probabilistic models, related methods also offer ways to circumnavigate the approximation of analytically intractable expectations. The key challenge in either setting is controlling the variance of gradient estimates: rec…
We develop a family of reformulations of an arbitrary consistent linear system into a stochastic problem. The reformulations are governed by two user-defined parameters: a positive definite matrix defining a norm, and an arbitrary discrete or continuous distribution over random matrices. Our reformulation has several e…
A new method solves l1-regularized optimization problems efficiently and sparsely.
problem l1-regularized optimization problems in machine learning.
method Orthant Based Proximal Stochastic Gradient Method (OBProx-SG)
result Promotes sparsity of solutions substantially and converges to global optimal solutions.
A parameter-free PGD algorithm for convex optimization.
problem Minimizing convex functions over convex sets.
method A fully adaptive AdaGrad variant of PGD without parameters or restarts.
result Optimal convergence rates for cumulative regret.
Large sectors of the recent optimization literature focused in the last decade on the development of optimal stochastic first order schemes for constrained convex models under progressively relaxed assumptions. Stochastic proximal point is an iterative scheme born from the adaptation of proximal point algorithm to nois…
New projection techniques reduce the frequency of projections in solving LCPs.
problem Solving linearly constrained problems efficiently with reduced projection frequency.
method Delayed projection technique to call a projection less frequently.
result Theoretical and practical improvements in convergence rates and efficiency.
New algorithm guarantees optimal convergence rate for stochastic optimization.
problem Optimal convergence rate for stochastic optimization algorithms.
method Regularized versions of Minimization by Incremental Surrogate Optimization (MISO) with arbitrary recurrent data sampling.
result Expected optimality gap converges at O(n−1/2) under general recurrent sampling schemes. Lower bounds on queries needed for finding stationary points in non-convex optimization.
problem Finding ε-stationary points in non-convex stochastic optimization. method Proving lower bounds on the number of queries required by stochastic first-order methods.
result Lower bounds on the number of queries required to find ε-stationary points are tight and optimal. A method for estimating the median of gradients in stochastic optimization.
problem Robust gradient estimation in stochastic optimization for various applications.
method Stochastic Proximal Point Method for median gradient estimation.
result The proposed method can converge even under heavy-tailed, state-dependent noise.
Estimates point counts on Riemannian varieties over finite fields.
problem Counting points on geometrically connected varieties over finite fields.
method Estimates point counts using Riemannian curvature and diameter.
result If sectional curvature and diameter grow, point counts diverge.