Gradient descent achieves exact linear convergence rate for symmetric matrix completion.
problem Low-rank symmetric matrix completion using gradient descent.
method Local analysis of gradient descent for symmetric matrices without additional assumptions.
result Closed-form expression of exact linear convergence rate matches practice.
The Sinkhorn-Knopp derivatives converge with linear rate.
problem Optimal transport problem with entropic regularization.
method Iterative proportional fitting procedure.
result Derivatives converge with linear rate.
Paper revisits set membership estimation for linear systems with relaxed disturbance bounds.
problem Set membership estimation for linear systems with disturbances bounded by convex sets.
method Adopted block-martingale small-ball condition and random perturbed control policies to establish convergence rates.
result Established convergence rates for disturbances bounded by general convex sets.
Paper analyzes TD( λ λ λ ) convergence rates for arbitrary features.
problem Convergence rates for linear TD( λ λ λ ) under arbitrary features. method Developed a novel stochastic approximation result for arbitrary features.
result Established L 2 L^2 L 2 convergence rates for linear TD( λ λ λ ) without linearly independent features assumption. New algorithm improves convergence rates for convex optimization problems.
problem Convex optimization problems with noisy stochastic data.
method Stochastic proximal point algorithm with weak linear regularity condition.
result Achieves $\mathcal{O}\left(\frac{1}{k}
ight)$ convergence rate for SPP.
Paper studies Gaussian approximation in linear regression with rates derived.
problem Gaussian approximation in online linear regression.
method Derives rates for constant learning rate settings, analyzes dependence on d d d and design matrix. result Rate of normal approximation is log n / n \sqrt{\log{n}/n} log n / n for large n n n . Active data collection improves convergence rates in operator learning.
problem Improving convergence rates in operator learning with linear target and stochastic input.
method Active data collection strategies with mean-zero stochastic process and continuous covariance kernels.
result Achieves arbitrarily fast error convergence rates with eigenvalue decay of covariance kernels.
We investigate finite-time decoupled convergence in nonlinear two-time-scale stochastic approximation.
problem Achieving decoupled convergence in nonlinear two-time-scale stochastic approximation.
method Nested local linearity assumption, suitable step size selection, convergence analysis of matrix cross term, fourth-order moment convergence rates.
result Finite-time decoupled convergence rates can be achieved in nonlinear two-time-scale stochastic approximation with proper step size selection.
Policy gradient methods achieve linear convergence in simple MDPs.
problem Analyzing convergence rates of policy gradient methods in finite MDPs.
method Connections with policy iteration to show linear convergence with large step-sizes.
result Policy gradient methods succeed with large step-sizes and achieve linear rate of convergence.
AdaLoss optimizes adaptive learning rates for efficient convergence in various models.
problem Efficiently optimizing adaptive learning rates for gradient descent methods.
method AdaLoss uses loss function information to dynamically adjust step sizes.
result AdaLoss achieves linear convergence in linear regression and robust global convergence in neural networks.
Stochastic gradient descent achieves polynomial convergence rates for noiseless linear models.
problem Convergence analysis of stochastic gradient descent in noiseless linear models.
method Fixed step-size stochastic gradient descent on least-square risk.
result Polynomial convergence rates depend on the regularities of the optimum and feature vectors.
The paper analyzes how hyperparameters affect SGD with momentum's convergence rate.
problem The role of hyperparameters in SGD with momentum's convergence rate.
method Theoretical analysis using a hyperparameters-dependent stochastic differential equation (hp-dependent SDE).
result The optimal linear rate of convergence depends on both the learning rate and the momentum coefficient.
Study identifies and validates a method for system identification of Markov jump linear systems.
problem System identification for autonomous Markov jump linear systems with complete state observations.
method Proposes switched least squares method for identification and derives rates of convergence.
result Data-independent rate of convergence is O ( log ( T ) / T ) \mathcal{O}\big(\sqrt{\log(T)/T} \big) O ( log ( T ) / T ) , showing strong consistency. Dropout speeds up convergence in shallow linear NNs, with a rate bound.
problem Analyzing convergence rate of Dropout in shallow linear NNs.
method Gradient flow analysis and Hessian examination.
result Bound on convergence rate depends on data, dropout probability, and NN width.
Gradient descent converges linearly for overparameterized linear networks.
problem Convergence of gradient descent for overparameterized neural networks.
method Local Polyak-Lojasiewicz and Descent Lemma for overparameterized linear models.
result Gradient descent achieves linear convergence for two-layer linear networks under relaxed assumptions.
Policy gradient converges linearly with Hadamard parameterization in tabular settings.
problem Convergence of policy gradient methods under Hadamard parameterization.
method Studied convergence rate and established linear convergence after k 0 k_0 k 0 iterations. result Algorithm converges linearly with rate $O(rac{1}{k})$ and faster locally after k 0 k_0 k 0 . Improved SGD for robust linear and ReLU regression with adversarial corruptions.
problem Robust regression with adversarial corruptions in streaming data.
method Stochastic gradient descent (SGD-exp) with exponentially decaying step size.
result Nearly linear convergence to true parameter with up to 50% Massart corruption rate.
Paper proposes a pre-conditioning technique to speed up gradient-descent convergence in distributed linear least-squares problems.
problem Expediting convergence of gradient-descent method for ill-conditioned distributed linear least-squares problems.
method Iterative pre-conditioning technique to improve convergence rate of gradient-descent method.
result Pre-conditioned gradient-descent achieves superlinear convergence for unique solutions and improved linear convergence otherwise.
Study shows a linear quadratic regulator's imitation learning converges globally.
problem Global convergence of imitation learning for linear quadratic regulators.
method Analyzed alternating gradient algorithm and established Q-linear rate of convergence.
result Established a unique saddle point for globally optimal policy and reward function.
Paper derives convergence rates and confidence intervals for LSA with Markovian noise.
problem Analyzing convergence rates and constructing confidence intervals for LSA with Markovian noise.
method Derives non-asymptotic Berry-Esseen bounds and multiplier block bootstrap procedure.
result Provides O ( n − 1 / 4 ) \mathcal{O}(n^{-1/4}) O ( n − 1/4 ) convergence rates and guarantees consistent inference. Linear Q-learning converges to a bounded set without divergence.
problem Proving linear Q-learning does not diverge and converges to a bounded set.
method No modifications to the original linear Q-learning algorithm, no Bellman completeness or near-optimality assumptions, only an ε-softmax behavior policy with adaptive temperature.
result First L 2 L^2 L 2 convergence rate of linear Q-learning iterates to a bounded set. This paper closes the gap on matching pursuit's convergence rate.
problem Improving the understanding of matching pursuit's convergence rate.
method Constructing a worst case dictionary to analyze matching pursuit's performance.
result Sharp characterization of matching pursuit's convergence rate as n − α n^{-α} n − α , with α ≈ 0.182 α \approx 0.182 α ≈ 0.182 . Guaranteed convergence for tensor factorization using Riemannian gradient descent.
problem Recovering tensor train format from linear measurements.
method Optimization over left-orthogonal TT format using Riemannian gradient descent on Stiefel manifold.
result RGD converges linearly to the ground-truth tensor with polynomial error growth in tensor order.
Unified framework for solving linear systems with improved convergence rates.
problem Efficiently solving linear systems with randomized batch-sampling methods.
method Developed a unified randomized batch-sampling Kaczmarz framework with concentration inequalities for analysis.
result Derived new expected linear convergence rate bounds that are tighter and more reflective of empirical behavior.
AM converges super-linearly for solving mixed linear regression problems.
problem Learning linear regressors from unlabeled observations in multiple linear regression models.
method Alternating Minimization (AM) algorithm, which alternates between label estimation and regression solving.
result AM converges super-linearly in certain parameter regimes, requiring only O(log log(1/ε)) iterations to achieve an error of ε.
SGD (Stochastic Gradient Descent) is a popular algorithm for large scale optimization problems due to its low iterative cost. However, SGD can not achieve linear convergence rate as FGD (Full Gradient Descent) because of the inherent gradient variance. To attack the problem, mini-batch SGD was proposed to get a trade-o…
Study derivative-free methods for linear policies in linear-quadratic systems.
problem Optimizing policies in linear-quadratic systems with limited derivative information.
method Derivative-free methods applied to linear policies over various noise and reward feedback settings.
result These methods converge to near-optimal policies with a polynomial number of zero-order evaluations.
Paper improves TD(0) convergence rate with LFA, i.i.d. samples, and averaging.
problem Improving convergence rate of TD(0) with linear function approximation.
method Polyak-Juditsky averaging, i.i.d. samples, strong mixing assumption.
result Established a new convergence rate for Mean-Square Error (MSE) of approximated function.
The paper analyzes how learning rate affects SGD and provides insights into optimal rates.
problem Understanding the impact of learning rate on stochastic gradient descent.
method Developed a learning-rate-dependent stochastic differential equation (lr-dependent SDE) to analyze SGD.
result Established a linear rate of convergence for SGD and found the optimal linear rate by analyzing the spectrum of the Witten-Laplacian.
Sharp bounds on weak convergence rate for rough volatility models.
problem Understanding the convergence rate in discretizing rough volatility models.
method Analyzing general and linear models to derive bounds.
result Sharper bound of \(H + 1/2\) for linear models.
Paper analyzes faster convergence rates for reinforcement learning from offline data.
problem Analyzing faster convergence rates for reinforcement learning from offline data.
method Fine analysis of reinforcement learning from offline data, providing fast rates for regret convergence.
result The paper provides fast rates for the regret convergence, showing that the level of exponentiation depends on the noise in the decision-making problem.
We show that Newton's method converges globally at a linear rate for objective functions whose Hessians are stable. This class of problems includes many functions which are not strongly convex, such as logistic regression. Our linear convergence result is (i) affine-invariant, and holds even if an (ii) approximate Hess…
Gradient descent converges linearly for deep linear networks under specific conditions.
problem Speed of convergence in gradient descent for deep linear neural networks.
method Analysis of gradient descent training for deep linear neural networks minimizing ℓ 2 \ell_2 ℓ 2 loss. result Gradient descent converges linearly under specific conditions on layer dimensions, initialization, and initial loss.
New algorithm combines gradient and coordinate descent steps for faster convergence.
problem Optimizing smooth convex functions over atom-spans.
method Blended matching pursuit combining coordinate descent and gradient descent.
result Derives linear convergence rates for non-strongly convex functions.
New algorithms prove fast convergence in complex min-max problems.
problem Proving fast convergence in nonconvex min-max optimization.
method Hamiltonian Gradient Descent (HGD) and Consensus Optimization (CO) algorithms.
result HGD and CO achieve linear convergence in various settings.
Paper shows TD learning without projection converges robustly.
problem Investigate convergence of TD learning with linear approx.
method Simple unprojected TD(0) with novel self-bounding property.
result TD(0) converges with rate O ~ ( 1 / T ) \widetilde{\mathcal{O}}(1/\sqrt{T}) O ( 1/ T ) . New algorithm for nonnegative tensor completion with linear convergence rate.
problem Tensor completion without known optimal sample complexity rate.
method Integer optimization using a specific 0-1 polytope gauge norm.
result Achieves information-theoretic rate with linear convergence.
Wide neural networks converge linearly to zero loss with feature learning.
problem Optimizing wide neural networks with feature learning guarantees.
method Gradient flow analysis for wide shallow and multi-layer NNs.
result Training loss converges linearly to zero for wide NNs under GF, demonstrating feature learning and better generalization.
PANDA improves linear discriminant analysis in high dimensions with minimal tuning.
problem Linear discriminant analysis in high-dimensional settings.
method PANDA: a tuning-insensitive method for linear discriminant analysis.
result PANDA achieves optimal convergence rates in estimation error and misclassification rate.
Study efficient derivative computation for nondifferentiable maps in machine learning.
problem Efficiently compute derivatives of fixed-point of nondifferentiable contractions.
method Iterative Differentiation (ITD), Approximate Implicit Differentiation (AID), and New Stochastic Implicit Differentiation (NSID).
result Established convergence rates for ITD, AID, and NSID, matching or improving smooth setting rates.
Study shows accelerated convergence of stochastic momentum methods in Wasserstein distances.
problem Performance sensitivity of momentum methods to noise in gradients.
method Stochastic momentum methods under a first-order stochastic oracle model.
result Linear convergence rates for AG and HB methods in Wasserstein metrics, robust to noise.
Unified approach to compute asymptotic constants using optimization.
problem Computing unknown constants in asymptotic expansions.
method Linear Least Squares and Tikhonov Linear Least Squares methods.
result Rigorous asymptotic estimates and convergence-rate guarantees.
Linear optimization is many times algorithmically simpler than non-linear convex optimization. Linear optimization over matroid polytopes, matching polytopes and path polytopes are example of problems for which we have simple and efficient combinatorial algorithms, but whose non-linear convex counterpart is harder and …
Stochastic gradient algorithms estimate the gradient based on only one or a few samples and enjoy low computational cost per iteration. They have been widely used in large-scale optimization problems. However, stochastic gradient algorithms are usually slow to converge and achieve sub-linear convergence rates, due to t…
We propose the stochastic average gradient (SAG) method for optimizing the sum of a finite number of smooth convex functions. Like stochastic gradient (SG) methods, the SAG method's iteration cost is independent of the number of terms in the sum. However, by incorporating a memory of previous gradient values the SAG me…
We analyze two novel randomized variants of the Frank-Wolfe (FW) or conditional gradient algorithm. While classical FW algorithms require solving a linear minimization problem over the domain at each iteration, the proposed method only requires to solve a linear minimization problem over a small \emph{subset} of the or…
This paper extends the convergence rate of DEQs with ReLU to any general activation.
problem Proving global convergence rate for DEQs with general activations.
method Developed a novel population Gram matrix and new form of dual activation with Hermite polynomial expansion.
result Gradient descent converges to a globally optimal solution at a linear rate for DEQs with general activations.
This paper proves AdaGrad and Adam converge linearly under PL inequality.
problem Understanding the convergence of adaptive gradient methods.
method Unified approach proving AdaGrad and Adam converge linearly under PL inequality.
result AdaGrad and Adam converge linearly when the cost function is smooth and satisfies PL inequality.