Generalizes smoothness conditions for optimization methods.
problem Optimization under non-uniform smoothness conditions.
method Develops a new analysis technique for bounding gradients.
result Obtains convergence rates for gradient descent and Nesterov's method.
This paper classifies solitons under specific tensor conditions.
problem Classifying solitons under vanishing conditions on the Weyl, Cotton, and Cao-Chen tensors.
method Analyzing complete conformal gradient solitons and using tensor conditions.
result Classification of complete nontrivial locally conformally flat conformal gradient solitons.
New bounds show BBVI's gradient variance matches SGD conditions, improving parameterization efficiency.
problem Understanding and improving the convergence of black-box variational inference (BBVI).
method Showed BBVI satisfies matching gradient variance bounds corresponding to the ABC condition for smooth and quadratically-growing log-likelihoods.
result Proven BBVI's gradient variance matches SGD conditions, with superior dimensional dependence for mean-field parameterization.
Faster reconstruction of compressed signals using conditional GAN and NPGD.
problem Recovering compressed signals from measurements.
method Network-based projected gradient descent (NPGD) combined with measurement-conditional generative adversarial networks (GANs/BEGANs).
result Significant speed-up in reconstruction (up to 140-175 times faster).
Unified algorithm for stochastic optimization with time-varying momentum converges under general conditions.
problem Optimizing functions with time-varying gradients and biases.
method Unified algorithm using a time-varying momentum term.
result Convergence of the unified algorithm under general conditions.
Proposes a new method for posterior sampling using MMD with negative distance kernel.
problem Posterior sampling and conditional generative modeling.
method Approximates joint distribution using discrete Wasserstein gradient flows of MMD with negative distance kernel.
result Establishes an error bound for posterior distributions and proves the method is a Wasserstein gradient flow.
Paper proposes a pre-conditioning technique to speed up gradient-descent convergence in distributed linear least-squares problems.
problem Expediting convergence of gradient-descent method for ill-conditioned distributed linear least-squares problems.
method Iterative pre-conditioning technique to improve convergence rate of gradient-descent method.
result Pre-conditioned gradient-descent achieves superlinear convergence for unique solutions and improved linear convergence otherwise.
Large deviations theory applied to policy gradient methods.
problem Understanding convergence of policy gradient methods in reinforcement learning.
method Large deviation rate function and contraction principle from large deviations theory.
result Convergence properties of policy gradient methods can be extended to various policy parametrizations.
The paper constructs optimal confidence bands for kernel gradient flow estimators.
problem Estimating generalization error and constructing confidence bands for kernel gradient flows.
method Established convergence rates and constructed optimal confidence bands under capacity-source condition.
result Optimal confidence bands for kernel gradient flows have shrinkage rates close to minimax optimal rates.
Paper proves SHB convergence with biased gradients and approximate step sizes.
problem Establishing convergence of SHB with biased gradients and approximate step sizes.
method Generalizes SHB convergence conditions for biased gradients, approximate step sizes, and block updating.
result Proves convergence of SHB with new conditions for biased gradients and approximate step sizes.
Paper proposes a policy gradient method for confounded POMDPs.
problem Estimating policy gradients for confounded POMDPs with continuous state and observation spaces.
method Developed a novel identification result to estimate policy gradients using offline data, solved conditional moment restrictions, and applied min-max learning with function approximation.
result Showed global convergence of the proposed algorithm in finding the optimal policy.
This work improves GAN stability with theoretical conditions.
problem GANs exhibit unstable behavior during training.
method Developed a theoretical framework and conditions for GAN stability.
result Constructs a GAN that fulfills stability conditions.
This paper analyzes adaptive gradient algorithms for better performance in ill-conditioned problems.
problem Poor performance of standard stochastic gradient algorithms in ill-conditioned problems.
method Non-asymptotic analysis of adaptive gradient algorithms (Adagrad and Stochastic Newton) for strongly convex objectives.
result Theoretical analysis and adaptation to practical applications like linear regression and regularized GLM.
Improved complexity for machine learning optimization methods.
problem Optimizing over-parametrized models in machine learning.
method Stochastic conditional gradient methods with interpolation-like conditions.
result Improved oracle complexities for finding optimal solutions.
The paper proves gradient estimates for nonlinear parabolic equations on smooth metric measure spaces.
problem Proving gradient estimates for nonlinear parabolic equations on smooth metric measure spaces.
method Using Souplet-Zhang type estimates and properties of Bakry-Emery Ricci tensor and weighted mean curvature.
result Gradient estimates for nonlinear parabolic equations on smooth metric measure spaces with Dirichlet boundary condition.
The objectives of this technical report is to provide additional results on the generalized conditional gradient methods introduced by Bredies et al. [BLM05]. Indeed , when the objective function is smooth, we provide a novel certificate of optimality and we show that the algorithm has a linear convergence rate. Applic…
We present Natural Gradient Boosting (NGBoost), an algorithm for generic probabilistic prediction via gradient boosting. Typical regression models return a point estimate, conditional on covariates, but probabilistic regression models output a full probability distribution over the outcome space, conditional on the cov…
We prove a gradient estimate for graphical spacelike mean curvature flow with a general Neumann boundary condition in dimension n=2. This then implies that the mean curvature flow exists for all time and converges to a translating solution.
New PG methods tackle nonconvex optimization with auto-conditioned stepsizes.
problem Optimizing nonconvex functions over convex sets.
method Auto-conditioned projected gradient (AC-PG) methods and stochastic variants.
result Achieved optimal iteration complexity for finding approximate stationary points.
The paper examines gradient ρ-Einstein solitons on specific manifolds and spacetimes.
problem Characterizing gradient ρ-Einstein solitons on doubly warped product manifolds.
method Analyzing necessary and sufficient conditions for doubly warped product manifolds to be gradient ρ-Einstein solitons, applying results to specific spacetime models.
result No 3-dimensional essentially conformally symmetric gradient ρ-Einstein soliton exists.
A new gradient boosting method improves interpretability of probabilistic models.
problem Learning interpretable yet accurate probabilistic models with limited rule complexity.
method A new objective function that measures the angle between risk gradient and condition output vector projection.
result Significantly improves comprehensibility/accuracy trade-off of fitted ensemble.
Study on gradient solitons on specific manifolds.
problem Existence and properties of gradient solitons on warped product manifolds.
method Analyzing necessary and sufficient conditions for the existence of generalized quasi Yamabe gradient solitons.
result Existence of non-trivial gradient Yamabe solitons on specific spacetimes.
Continuous-time SGD converges under certain conditions, useful for deep learning.
problem Minimizing population expected loss in learning problems.
method Continuous-time approximation of stochastic gradient descent.
result Establishes sufficient conditions for convergence, applicable to overparametrized neural networks.
WaveGrad generates high-fidelity audio using gradient estimation.
problem Generating high-fidelity audio efficiently.
method Conditional model using score matching and diffusion models, iteratively refining a Gaussian white noise signal.
result WaveGrad can generate high-fidelity audio samples using as few as six iterations.
New algorithm tackles nonconvex machine learning problems with adaptive normalization and independent sampling.
problem Nonconvex machine learning problems with generalized-smoothness.
method Adaptive gradient normalization, independent sampling, and gradient clipping.
result Achieves an O(ε^(-4)) sample complexity for fast convergence.
A new biased gradient descent method for conditional stochastic optimization.
problem Challenges in constructing unbiased gradient estimators for conditional stochastic optimization.
method Proposes a biased stochastic gradient descent (BSGD) algorithm and analyzes its sample complexities.
result Establishes sample complexities of BSGD for various objectives and shows that BSpiderBoost matches the lower bound complexity.
In 1963, Polyak proposed a simple condition that is sufficient to show a global linear convergence rate for gradient descent. This condition is a special case of the Łojasiewicz inequality proposed in the same year, and it does not require strong convexity (or even convexity). In this work, we show that this much-older…
Gradient descent on MMD GAN parameter space converges globally to target distribution.
problem Convergence of gradient descent in Maximum Mean Discrepancy (MMD) GANs.
method Proposes a parametric kernelized gradient flow that mimics the min-max game in gradient regularized MMD GAN.
result Gradient descent on the generator's parameter space in gradient regularized MMD GAN is globally convergent to the target distribution under certain conditions.
Given a convex optimization problem and its dual, there are many possible first-order algorithms. In this paper, we show the equivalence between mirror descent algorithms and algorithms generalizing the conditional gradient method. This is done through convex duality, and implies notably that for certain problems, such…
New approach proves convergence of SA and SGD with weaker conditions.
problem Proving convergence of SA and SGD with relaxed noise conditions.
method Introduces GSLLN to decouple function and noise properties.
result Derives sufficient conditions for convergence of SA and SGD.
New analysis shows GMD can converge linearly under PL-like conditions.
problem Establishing linear convergence for generalized mirror descent.
method PL-based analysis for time-dependent mirrors, Taylor-series approach for stochastic GMD.
result Linear convergence of stochastic GMD under PL-like conditions.
The paper investigates geometrical aspects of static spacetime with almost gradient Ricci solitons.
problem Geometrical properties of static spacetime with almost gradient Ricci solitons.
method Analyzing conditions and properties of static spacetime with almost gradient Ricci solitons.
result Conditions and properties of static spacetime with almost gradient Ricci solitons are determined.
We consider a condition on the Ricci curvature involving vector fields, which is broader than the Bakry-Émery Ricci condition. Under this condition volume comparison, Laplacian comparison, isoperimetric inequality and gradient bounds are proven on the manifold. Specializing to the Bakry-Émery Ricci curvature condition,…
Polyak step size GD reaches final radius of convergence after log iterations.
problem Statistical and computational complexities of Polyak step size GD.
method Generalized smoothness and Lojasiewicz conditions, stability of gradients.
result Polyak step size GD reaches final statistical radius of convergence after logarithmic number of iterations.
The natural gradient of ELBO vanishes in unconstrained optimization, simplifying learning.
problem The gap between evidence and ELBO has a vanishing natural gradient.
method Analyzes the Fisher-Rao gradient of ELBO and its implications for learning.
result Maximizing ELBO is equivalent to minimizing KL divergence, simplifying learning.
Zero loss is achievable in overparametrized DL networks under specific conditions.
problem Achieving zero loss in overparametrized deep learning networks.
method Determine sufficient conditions for zero loss attainability and present an explicit construction of zero loss minimizers.
result Explicit minimizers for zero loss in overparametrized DL networks are constructed without gradient descent.
ScaledGD improves gradient descent for ill-conditioned low-rank matrix estimation.
problem Efficiently solving ill-conditioned low-rank matrix estimation problems.
method Scaled Gradient Descent (ScaledGD) with adaptive pre-conditioners.
result Linear convergence rate independent of condition number, low per-iteration cost.
New algorithms improve distributed optimization under mild variance conditions.
problem Improving distributed optimization for large-scale machine learning problems.
method Revisited Federated Averaging and SCAFFOLD algorithms under a general variance condition.
result Established convergence results for smooth nonconvex objective functions under mild variance conditions.
Unified probabilistic gradient boosting for entire conditional distribution modeling.
problem Creating accurate probabilistic forecasts from regression tasks.
method Unified probabilistic gradient boosting framework using XGBoost and LightGBM, modeling conditional moments or CDF via Normalizing Flows.
result Achieves state-of-the-art forecast accuracy.
New method tackles nonconvex-nonconcave problems with local KL condition.
problem Nonconvex-nonconcave minimax problems under varying KL conditions.
method Inexact proximal gradient method for KL-structured subproblems.
result Complexity guarantees for approximate stationary points.
In this paper we classify the four dimensional gradient shrinking solitons under certain curvature conditions satisfied by all solitons arising from finite time singularities of Ricci flow on compact four manifolds with positive isotropic curvature. As a corollary we generalize a result of Perelman on three dimensional…
The paper establishes gradient estimates for harmonic and heat equation solutions on manifolds with boundary.
problem Gradient estimates for harmonic and heat equation solutions on manifolds with boundary.
method Yau and Souplet-Zhang type gradient estimates for harmonic and heat equation solutions under Dirichlet boundary condition.
result Established gradient estimates for harmonic and heat equation solutions on manifolds with boundary.
Study on shrinking solitons of generalized Ricci flow.
problem Characterizing shrinking solitons in generalized Ricci flow.
method Analyzing gradient shrinking solitons and pluriclosed solitons on compact manifolds.
result First non-trivial shrinking generalized soliton constructed.
Characterizes gradient Yamabe solitons with specific conditions.
problem Understanding properties of gradient Yamabe solitons.
method Proved conditions leading to constant scalar curvature, subharmonicity, and harmonic potential.
result Gradient Yamabe solitons under certain conditions are of constant scalar curvature.
New unbiased gradient estimators for complex optimization problems.
problem Unbiased and variance-limited gradient estimation for conditional stochastic optimization.
method Developed multilevel Monte Carlo gradient estimators for conditional stochastic optimization problems.
result Unbiased and finite variance gradient estimators for conditional stochastic optimization problems.
The study finds a lower bound for the diameter of gradient ρ-Einstein solitons.
problem Estimating the diameter of gradient ρ-Einstein solitons.
method Using mathematical conditions and properties of solitons to derive a lower bound.
result A lower bound for the diameter of gradient ρ-Einstein solitons is established.
The study characterizes GRW spacetimes with gradient solitons and phantom era.
problem Characterizing generalized Robertson-Walker spacetimes with gradient solitons.
method Examined gradient type Ricci solitons and (m,τ)-quasi Einstein solitons in GRW spacetimes. result Demonstrated that GRW spacetimes can be Robertson-Walker or phantom era spacetimes under certain conditions.
ScaledGD accelerates ill-conditioned low-rank estimation.
problem Slow convergence of gradient descent in ill-conditioned problems.
method Scaled gradient descent (ScaledGD) with preconditioning.
result Linear convergence rate independent of condition number.