New bound relaxes uniform gradient norm assumptions for PAC-Bayesian bounds.
problem Generalization bounds with strict assumptions like uniformly bounded loss.
method Relax uniform bounds assumptions to on-average bounded loss and gradient norm.
result Proposes a new generalization bound with a surrogate of model complexity.
Gradient bounds and Liouville theorems for quasi-linear equations on manifolds with nonnegative Ricci curvature.
problem Establishing bounds and theorems for solutions to quasi-linear elliptic equations on compact manifolds with nonnegative Ricci curvature.
method Gradient bounds, Liouville-type theorems, local splitting theorem, Harnack-type inequality, ABP estimate.
result Gradient bounds and Liouville-type theorems for solutions to quasi-linear equations on compact manifolds with nonnegative Ricci curvature.
The study classifies specific types of solitons with bounded scalar curvature.
problem Classifying quasi-Yamabe gradient solitons with bounded scalar curvature.
method Analyzing complete, nontrivial solitons with scalar curvature bounded above or below.
result Classification of specific types of solitons with bounded scalar curvature.
The study finds a lower bound for the diameter of gradient ρ-Einstein solitons.
problem Estimating the diameter of gradient ρ-Einstein solitons.
method Using mathematical conditions and properties of solitons to derive a lower bound.
result A lower bound for the diameter of gradient ρ-Einstein solitons is established.
The paper provides gradient estimates for solutions on manifolds with integral Ricci bounds.
problem Global regularity estimates for solutions of Δ u = f Δu = f Δ u = f on Riemannian manifolds. method Proves L p L^p L p -gradient estimates under integral Ricci bounds and constructs a counterexample. result Optimal constant lower bounds on Ricci curvature are shown in the pointwise sense.
Paper establishes a generalization bound for gradient flow using a data-dependent kernel.
problem Understanding the generalization properties of gradient-based optimization methods.
method Establishes a generalization bound for gradient flow through a data-dependent kernel called the loss path kernel (LPK).
result The LPK captures the entire training trajectory and leads to tighter generalization guarantees.
Lower bounds on queries needed for finding stationary points in non-convex optimization.
problem Finding ε ε ε -stationary points in non-convex stochastic optimization. method Proving lower bounds on the number of queries required by stochastic first-order methods.
result Lower bounds on the number of queries required to find ε ε ε -stationary points are tight and optimal. Generalizes smoothness conditions for optimization methods.
problem Optimization under non-uniform smoothness conditions.
method Develops a new analysis technique for bounding gradients.
result Obtains convergence rates for gradient descent and Nesterov's method.
Sharp Gaussian bounds derived for Schrödinger kernel on Ricci solitons.
problem Analyzing Schrödinger heat kernel on gradient shrinking Ricci solitons.
method Deriving sharp Gaussian upper bounds for the Schrödinger heat kernel.
result Sharp upper and lower bounds for eigenvalues of the Schrödinger operator.
Derives gradient bounds for f-heat equations on manifolds with Bakry-Emery Ricci curvature.
problem Gradient estimates for positive solutions of f-heat equations on manifolds with specific curvature conditions.
method Applies Li-Yau gradient estimates to positive solutions of the f-heat equation on closed manifolds with Bakry-Emery Ricci curvature bounded below.
result Derives Li-Yau gradient bounds for positive solutions of the f-heat equation.
Proves convergence of gradient Ricci shrinkers with uniform bounds.
problem Compactness and energy concentration in gradient Ricci shrinkers.
method Bubble-tree convergence and local energy analysis.
result No energy concentrates in neck regions, leading to a local diffeomorphism finiteness theorem.
5D shrinking solitons with bounded curvature are rigid.
problem Characterizing 5D shrinking gradient Ricci solitons.
method Proving rigidity for solitons with bounded curvature.
result 5D shrinking gradient Ricci solitons with bounded curvature are rigid.
Characterizes Kähler-hyperbolicity of bounded symmetric domains based on rank and genus.
problem Understanding the Kähler-hyperbolicity of bounded symmetric domains.
method Defines Kähler-hyperbolicity length by rank and genus, and characterizes it through a special Bergman potential.
result Establishes a unique constant for Kähler-hyperbolicity based on gradient length of a Bergman potential.
Lower bounds show many sampling algorithms need many gradient queries.
problem Sampling from strongly log-concave densities in high dimensions.
method Information theory and stochastic gradient methods.
result Lower bound on number of gradient queries needed.
In this paper, we prove the compactness theorem for gradient Ricci solitons. Let ( M α , g α ) (M_α, g_α) ( M α , g α ) be a sequence of compact gradient Ricci solitons of dimension n ≥ 4 n\geq 4 n ≥ 4 , whose curvatures have uniformly bounded L n 2 L^{\frac{n}{2}} L 2 n norms, whose Ricci curvatures are uniformly bounded from below with uniformly lower bounded vol…
In this work we revisit gradient regularization for adversarial robustness with some new ingredients. First, we derive new per-image theoretical robustness bounds based on local gradient information. These bounds strongly motivate input gradient regularization. Second, we implement a scaleable version of input gradient…
New decay estimates for scalar curvature of steady gradient Ricci solitons.
problem Understanding scalar curvature behavior in steady gradient Ricci solitons.
method Using μ-bubbles introduced by Gromov.
result Provide new decay estimates for scalar curvatures.
New bounds for model generalization under deterministic gradient descent.
problem Establishing generalization bounds for models trained with gradient descent methods.
method PAC-Bayesian bounds for deterministic optimisation algorithms.
result Fully computable bounds that depend on initial distribution and Hessian.
Improved bounds for proximal gradient algorithms with computational errors.
problem Analyzing convergence of proximal gradient algorithms with inaccuracies.
method Deriving new tighter deterministic and probabilistic bounds for convex composite problems.
result Probabilistic bounds are more robust and accurate for algorithm verification and performance guarantees.
Short note on soft-max and policy gradients in bandit problems using Lyapunov functions.
problem Analyzing soft-max and policy gradient methods in bandit problems.
method Lyapunov function argument for soft-max and differential equations for policy gradient algorithms.
result Regret bounds for soft-max and a different policy gradient algorithm in bandit problems.
Improved Gaussian process regression with tighter log marginal likelihood bounds.
problem Improving predictive performance in Gaussian process regression models.
method Lower bound on log marginal likelihood using conjugate gradients.
result Improved predictive performance compared to other conjugate gradient based approaches.
Paper relaxes stability and generalization assumptions for SGD.
problem Stability and generalization for SGD under restrictive assumptions.
method Introduces on-average model stability and develops novel bounds.
result First-ever-known fast bounds in low-noise setting using stability approach.
New research shows existing information-theoretic methods can't establish minimax rates for gradient descent in stochastic convex optimization.
problem Establishing minimax rates for gradient descent in stochastic convex optimization using information-theoretic methods.
method Examined several information-theoretic frameworks including input-output mutual information bounds, conditional mutual information bounds, PAC-Bayes bounds, and their variants.
result Proved that none of the examined information-theoretic frameworks can establish minimax rates for gradient descent in stochastic convex optimization.
Unified bounds for random subset generalization error and improved SGD Langevin dynamics.
problem Generalization error bounds for random subsets and stochastic gradient Langevin dynamics.
method Unified framework based on Hellström and Durisi's work, extending bounds for Langevin dynamics.
result Unified and refined bounds for generalization error in stochastic gradient Langevin dynamics.
New oracles improve stochastic optimization with noisy or biased measurements.
problem Optimizing functions with noisy or biased measurements.
method Introduced biased gradient oracles for stochastic optimization, analyzed RSG and SGD algorithms with these oracles.
result Derived non-asymptotic bounds for convergence rates of algorithms with biased gradient oracles.
The natural gradient of ELBO vanishes in unconstrained optimization, simplifying learning.
problem The gap between evidence and ELBO has a vanishing natural gradient.
method Analyzes the Fisher-Rao gradient of ELBO and its implications for learning.
result Maximizing ELBO is equivalent to minimizing KL divergence, simplifying learning.
In this very short note we prove a lower bound for the scalar curvature of certain steady gradient Ricci solitons.
Maxout networks study gradients and propose initialization strategies.
problem Complexity in input-output Jacobian distribution complicates stable parameter initialization.
method Obtained bounds on moments of gradients and formulated initialization strategies.
result Parameter initialization strategies improve training of deep maxout networks.
Gradient descent fails to learn simple neural networks efficiently.
problem Learning one-layer neural networks efficiently using gradient descent.
method Gradient descent and statistical query algorithms.
result Superpolynomial lower bounds for learning one-layer neural networks.
Improved DP algorithms for non-convex optimization with tighter generalization bounds.
problem Private stochastic non-convex optimization in high-dimensional spaces.
method Differential privacy techniques, including adaptive algorithms like DP RMSProp and DP Adam, combined with adaptive data analysis.
result Achieved a sharper rate of p 4 / n \sqrt[4]{p}/\sqrt{n} 4 p / n for population loss, improving upon previous bounds. Paper formalizes and analyzes a new bound for variational inference.
problem Lack of theoretical guarantees in variational algorithms.
method Introduces VR-IWAE bound, a generalization of IWAE.
result VR-IWAE bound leads to unbiased gradient estimators.
New findings show margins are not sufficient for explaining gradient boosting performance.
problem The inadequacy of margin explanations in explaining the performance of gradient boosting.
method Demonstrated and proved a stronger margin-based generalization bound for boosted classifiers.
result Proved a stronger margin-based generalization bound that explains the performance of modern gradient boosters.
SGD handles label noise with bounds improving over SGLD.
problem Label noise in non-convex optimization.
method Stochastic gradient descent with uniform dissipativity and smoothness conditions, using Wasserstein distance and algorithmic stability.
result Generalization error bounds with a rate of n − 2 / 3 n^{-2/3} n − 2/3 , better than SGLD's n − 1 / 2 n^{-1/2} n − 1/2 . New algorithm guarantees performance on noisy data.
problem Learning with noisy data and heavy-tailed distributions.
method Anytime online-to-batch conversion for smooth objectives.
result Stochastic gradient-based algorithm with sub-Gaussian error bounds.
Many continuous control tasks have bounded action spaces. When policy gradient methods are applied to such tasks, out-of-bound actions need to be clipped before execution, while policies are usually optimized as if the actions are not clipped. We propose a policy gradient estimator that exploits the knowledge of action…
New bounds show BBVI's gradient variance matches SGD conditions, improving parameterization efficiency.
problem Understanding and improving the convergence of black-box variational inference (BBVI).
method Showed BBVI satisfies matching gradient variance bounds corresponding to the ABC condition for smooth and quadratically-growing log-likelihoods.
result Proven BBVI's gradient variance matches SGD conditions, with superior dimensional dependence for mean-field parameterization.
Optimizes shortfall risk using gradient-based methods.
problem Optimizing utility-based shortfall risk measures.
method Gradient-based stochastic optimization, non-asymptotic bounds derivation.
result Non-asymptotic convergence rate for optimizing UBSR.
Paper extends Aronson-Bénilan estimates for porous medium equations on manifolds with negative curvature.
problem Estimating gradients for porous medium equations on manifolds with negative curvature.
method Develops Aronson-Bénilan gradient estimates for porous medium equations under lower bounds of N N N -weighted Ricci curvature with N < 0 N < 0 N < 0 . result Generalizes gradient estimates for porous medium equations to manifolds with negative curvature.
Researchers estimate gradients of solutions to a Finslerian Allen-Cahn equation.
problem Estimating gradients of solutions to a specific type of partial differential equation.
method Using the Finslerian Allen-Cahn equation as an Euler-Lagrange equation to a Liapunov entropy functional, proving gradient estimates on compact and noncompact Finsler metric measure spaces.
result Global and local gradient estimates of positive solutions to the Finslerian Allen-Cahn equation.
Paper develops probabilistic bounds for a stochastic gradient algorithm in non-convex problems.
problem Stochastic optimization in non-convex finite sum problems.
method Develops a new dimension-free Azuma-Hoeffding type bound for a martingale difference sequence.
result Empirical results show superior probabilistic performance of Prob-SARAH compared to other algorithms.
SGD generalization bounds derived from information theory.
problem Understanding generalization of SGD for non-convex functions.
method Combining information-theoretic bounds with perturbation analysis.
result Upper bounds on SGD's generalization error based on gradient variance and function smoothness.
Sharp Lipschitz bounds and gradient estimates for fully nonlinear parabolic equations.
problem Understanding moduli of continuity for fully nonlinear parabolic equations.
method Proving moduli of continuity of viscosity solutions are subsolutions of one-dimensional parabolic equations.
result Sharp Lipschitz bounds and gradient estimates for fully nonlinear parabolic equations with bounded initial data.
This work bounds the run-time of nonconvex optimization with early stopping.
problem Bounding the expected run-time of nonconvex optimization with early stopping.
method Derives conditions for well-defined early stopping based on validation function norms and bounds the expected number of iterations and gradient evaluations.
result Guarantees the validity of early stopping and provides bounds on the expected run-time for various optimization algorithms.
The paper proves gradient estimates for a weighted p-Laplacian equation on Riemannian manifolds.
problem Gradient estimates for a weighted p-Laplacian equation on Riemannian manifolds.
method Assumes a Sobolev inequality and integral Ricci bounds, proving local gradient estimates and Liouville type results.
result Proves local gradient estimates and Liouville type results on manifolds with lower bounds of Ricci curvature.
DIFF2 improves differential privacy in nonconvex optimization with better utility bounds.
problem Improving differential privacy in nonconvex optimization with better utility bounds.
method DIFF2 constructs a differential private global gradient estimator using gradient differences.
result DIFF2 achieves a utility of \(\widetilde O(d^{2/3}/(n\varepsilon_{\mathrm{DP}})^{4/3})\), significantly better than \(\widetilde O(\sqrt{d}/(n\varepsilon_{\mathrm{DP}}))\).
Paper investigates conditions for independence of weak gradients on metric spaces.
problem Dependence of weak gradients on p p p in arbitrary metric measure spaces. method Investigates the Bounded Interpolation Property to ensure independence of weak gradients.
result Bounded Interpolation Property guarantees independence of weak gradients.
The paper studies steady solitons with curvature decay and proves their smoothness.
problem Analyzing the properties of steady solitons with curvature decay.
method Bootstrap regularity in harmonic coordinates using the soliton equation.
result Steady gradient Ricci solitons are asymptotically cylindrical under certain curvature decay conditions.
Study 4D solitons with specific curvature properties, proving curvature bounds and classifying solutions.
problem Investigate 4D gradient solitons with specific curvature properties.
method Analyze 4D gradient steady and shrinking solitons with nonnegative or half nonnegative isotropic curvature.
result Prove 2-nonnegativity of Ricci curvature and bound the curvature tensor for ancient solutions.