New methods using natural gradient for structured optimization.
problem Structured optimization problems.
method Structured second-order methods via natural gradient descent.
result Efficiency demonstrated on non-convex and deep learning problems.
Paper examines the structure of stochastic gradients in deep learning.
problem Exploring the structure and heavy tails of stochastic gradients in deep learning.
method Conducted formal statistical tests on stochastic gradients and gradient noise.
result Stochastic gradients and gradient noise do not exhibit power-law heavy tails, but their covariance spectra do.
NES optimizes discrete structured VAEs effectively without gradient propagation.
problem Learning high-dimensional discrete latent spaces in generative models.
method Natural Evolution Strategies (NES) for gradient-free optimization of discrete structures.
result NES effectively optimizes discrete structured VAEs, comparable to gradient-based methods.
The study classifies and disproves gradient properties of certain solitons on specific Lie groups.
problem Characterizing and proving non-graduation of solitons on specific Lie groups.
method Proving structure theorems and analyzing specific examples of solitons.
result Examples of solitons that cannot be made gradient, including specific Lie groups.
Gradient flow converges to a minimal convex structure.
problem Finding the minimal convex structure in hyperbolic manifolds.
method Weil-Petersson gradient vector field of renormalized volume.
result The flow converges to the structure with minimum convex core volume.
Constructs explicit solutions to Spin(7)-structures gradient flow.
problem Finding explicit solutions to Spin(7)-structures gradient flow.
method Expressed Spin(7)-torsion tensor and gradient flow in terms of torsion forms; used these formulae to find solutions.
result Found explicit solutions including a shrinking soliton on SU(3) and another on a T7-bundle over S1. StructureBoost improves gradient boosting for complex categorical variables efficiently.
problem Efficiently handling complex categorical variables with known structure.
method Two methods to overcome computational obstacles in SCDT enumeration for structured categorical variables.
result StructureBoost outperforms existing packages on complex categorical problems.
Gradient flow studies Spin(7)-structures on compact 8-manifolds.
problem Formulating and studying the gradient flow of Spin(7)-structures.
method Negative gradient flow of an energy functional of Spin(7)-structures.
result Short-time existence and uniqueness of solutions to the flow.
The paper studies Einstein-type structures in warped product manifolds.
problem Characterizing Einstein-type structures in warped product manifolds.
method Analyzing conditions for minimal, totally umbilical, and geodesic immersions.
result Characterization of rotational hypersurfaces in RimesfRn. Study on Yamabe solitons with applications and structure elucidation.
problem Understanding the structure of Yamabe solitons and their applications.
method Investigation of complete gradient conformal solitons under specific conditions.
result Affirmative partial answer to Yamabe soliton conjecture.
Mean curvature flow is not a gradient flow on two nondegenerate metric spaces.
problem Whether mean curvature flow is a gradient flow on nondegenerate metric spaces of simple closed plane curves.
method Examined two nondegenerate metric spaces: uniformness-preserving and curvature-weighted structures.
result Mean curvature flow is not a gradient flow on either metric space.
This work explores gradient flows and Riemannian structure in Gromov-Wasserstein geometry for data with global structure.
problem Suitable geometry for tasks requiring preservation of global data structure.
method Study of gradient flows and Riemannian structure in Gromov-Wasserstein geometry for distributions on \(\mathbb{R}^d\).
result Established a Benamou-Brenier-like formula for IGW and derived the IGW gradient.
We show that every complete nontrivial gradient Yamabe soliton admits a special global warped product structure with a one-dimensional base. Based on this, we prove a general classification theorem for complete nontrivial locally conformally flat gradient Yamabe solitons.
Uniqueness proven for specific types of geometric structures.
problem Proving uniqueness of asymptotically conical gradient shrinking solitons.
method Extends Kotschwar and Wang's argument for uniqueness of AC gradient shrinking Ricci solitons.
result G_2-structures are equivalent if asymptotically conical and asymptotic to the same closed G_2-cone.
A gradient-based method learns the structure of TAN for Bayesian network classifiers.
problem Learning the structure of Bayesian networks is difficult.
method A distribution over graph structures learned via gradient-based optimization.
result Consistently outperforms random and Chow-Liu TAN structures.
Extends geometric structures to manifolds with new operators.
problem No specific problem stated; extending geometric structures.
method Defines new gradient and Laplace operators on manifolds with geometric structures.
result Provides properties of the new operators.
Transformers learn causal structure through gradient descent on self-attention mechanisms.
problem Understanding how transformers learn causal structure during training.
method In-context learning task and simplified two-layer transformer model.
result Gradient descent on a simplified transformer learns to encode latent causal graphs.
Research describes all possible gradient vector fields on a sphere with up to ten singular points.
problem Characterizing gradient vector fields on a sphere with limited singular points.
method Using a graph to represent one-dimensional stable manifolds, specifying singularities and connections.
result Identified all topological structures of codimension one gradient vector fields on a sphere with up to ten singular points.
Study proves structure results for homogeneous spaces supporting specific equations.
problem Proving structure results for homogeneous spaces supporting specific equations.
method Analyzing homogeneous spaces with non-constant solutions to two general classes of equations involving the Hessian and an invariant 2-tensor.
result Generalizes rigidity results for gradient Ricci solitons and warped product Einstein metrics.
Causal structure learning has been a challenging task in the past decades and several mainstream approaches such as constraint- and score-based methods have been studied with theoretical guarantees. Recently, a new approach has transformed the combinatorial structure learning problem into a continuous one and then solv…
Introduces new info-geometric structure for dynamics on graphs and hypergraphs.
problem Modeling dynamics on discrete structures like graphs and hypergraphs.
method Introduces two dually flat structures: one on vertex space and another on edge space.
result Extends gradient flows to include nonequilibrium dynamics.
We study the geometry at infinity of expanding gradient Ricci solitons of dimension greater than two with finite asymptotic curvature ratio without curvature sign assumptions. We mainly prove that they have a cone structure at infinity.
This work investigates how gradient-based learning performs with structured data, revealing issues and improvements.
problem Gradient-based learning under structured data, particularly with a spiked covariance structure.
method Investigates the effect of a spiked covariance structure on gradient-based feature learning and proposes weight normalization.
result Gradient-based dynamics may fail to recover the true direction in anisotropic settings, but weight normalization can improve performance.
We prove that a nontrivial complete generalized quasi Yamabe gradient soliton (M; g) must be a quasi Yamabe gradient soliton on each connected component of M and that a nontrivial complete locally conformally at generalized quasi Yamabe gradient soliton has a special warped product structure.
We propose a novel data-dependent structured gradient regularizer to increase the robustness of neural networks vis-a-vis adversarial perturbations. Our regularizer can be derived as a controlled approximation from first principles, leveraging the fundamental link between training with noise and regularization. It adds…
In this paper we extend some well-known rigidity results for conformal changes of Einstein metrics to the class of generalized quasi-Einstein (GQE) metrics, which includes gradient Ricci solitons. In order to do so, we introduce the notions of conformal diffeomorphisms and vector fields that preserve a GQE structure. W…
New method improves structure learning on sparse graphs.
problem Structure learning on sparse directed acyclic graphs (DAGs).
method Bregman proximal gradient method to address non-convex, high-curvature problem.
result Significantly improved convergence and efficiency.
These notes aim to shed light on the recently proposed structured projected intermediate gradient optimization technique (SPIGOT, Peng et al., 2018). SPIGOT is a variant of the straight-through estimator (Bengio et al., 2013) which bypasses gradients of the argmax function by back-propagating a surrogate "gradient." We…
Study shows Kähler gradient Ricci solitons have limited symmetry groups.
problem Understanding the symmetry groups of Kähler gradient Ricci solitons.
method Connection to almost contact metric structure.
result Group of isometries is at most n^2, with equality characterized.
Evolution Strategies (ES) are a powerful class of blackbox optimization techniques that recently became a competitive alternative to state-of-the-art policy gradient (PG) algorithms for reinforcement learning (RL). We propose a new method for improving accuracy of the ES algorithms, that as opposed to recent approaches…
We develop a more efficient NGD method for structured parameters.
problem Computational challenges in NGD for structured parameter spaces.
method Local-parameter coordinates to simplify Fisher-matrix computations.
result New structured second-order algorithms and learning methods.
Gradient flows of neural networks converge to optimal values or diverge, with thresholds and asymptotic behaviors.
problem Understanding the convergence and divergence of gradient flows in neural networks.
method Analysis of gradient flows on loss landscapes of neural networks using o-minimal structures.
result Gradient flows either converge to optimal values or diverge to infinity, with thresholds and asymptotic behaviors.
GIT uses gradient estimators to target interventions for causal discovery.
problem Challenges in inferring causal structure from observational data.
method GIT uses gradient estimators to target interventions for causal discovery.
result GIT performs on par with competitive baselines, surpassing them in low-data regimes.
PolarGrad optimizes deep learning models by considering matrix structure, outperforming Adam and Muon.
problem Efficient optimization of large-scale neural networks and language models.
method A unifying framework for analyzing matrix-aware preconditioned methods, including PolarGrad.
result PolarGrad outperforms Adam and Muon in various tasks.
Unified framework for gradient estimation in combinatorial spaces.
problem Scaling relaxed gradient estimators to large combinatorial distributions.
method Introducing stochastic softmax tricks within the perturbation model framework.
result Stochastic softmax tricks improve model performance and discover more latent structure.
We describe the structure of the Ricci tensor on a locally homogeneous Lorentzian gradient Ricci soliton. In the non-steady case, we show the soliton is rigid in dimensions three and four. In the steady case, we give a complete classification in dimension three.
Gradient descent finds a global minimum in training deep neural networks despite the objective function being non-convex. The current paper proves gradient descent achieves zero training loss in polynomial time for a deep over-parameterized neural network with residual connections (ResNet). Our analysis relies on the p…
Gradient descent on MMD GAN parameter space converges globally to target distribution.
problem Convergence of gradient descent in Maximum Mean Discrepancy (MMD) GANs.
method Proposes a parametric kernelized gradient flow that mimics the min-max game in gradient regularized MMD GAN.
result Gradient descent on the generator's parameter space in gradient regularized MMD GAN is globally convergent to the target distribution under certain conditions.
Gradient guidance improves diffusion models for optimizing specific objectives.
problem Improving diffusion models for specific optimization tasks.
method Established a mathematical framework for gradient-guided diffusion, linking it to optimization theory. Developed a modified gradient guidance method and iteratively fine-tuned version.
result Gradient-guided diffusion models are essentially solutions to regularized optimization problems, preserving latent structure.
Study describes bifurcations of gradient flows on 2-sphere with holes.
problem Analyzing gradient flows on a 2-sphere with up to six singular points.
method Using separatrix diagrams to specify saddle-node and saddle connections.
result Identified all possible topological structures of bifurcations.
Gradient descent implicitly favors group sparsity in neural networks.
problem Understanding implicit regularization in neural networks for structured sparsity.
method Novel neural reparameterization for diagonally grouped linear networks.
result Gradient descent without explicit regularization biases towards group sparsity.
Embeds sparsity in deep neural networks, allowing exact zero parameters.
problem Learning sparse structures in deep networks.
method Embeds sparsity into neural network structure, allowing exact zero parameters during training.
result Can learn both structured and unstructured sparsity.
New research shows input-gradients can be manipulated without changing model's core function, challenging their use for model interpretation.
problem Current methods for model interpretability using input-gradients are flawed due to their arbitrary manipulability.
method Investigated by reinterpreting logits as unnormalized log-densities, proposing novel approximations for score-matching.
result Improving alignment between implicit density model and data distribution enhances gradient structure and explanatory power.
The gradient noise of SGD is considered to play a central role in the observed strong generalization abilities of deep learning. While past studies confirm that the magnitude and the covariance structure of gradient noise are critical for regularization, it remains unclear whether or not the class of noise distribution…
Paper analyzes convergence of dynamic policy gradient for MDPs, improving performance in finite-time problems.
problem Optimal policies in finite-time MDPs are not stationary and require epoch-specific training.
method Introduces dynamic policy gradient combining dynamic programming and policy gradient, analyzes convergence for softmax parametrisation.
result Dynamic policy gradient training exploits finite-time structure, leading to better convergence bounds.
Study on dimensions of Killing vector fields on gradient Ricci solitons.
problem Estimating dimensions of Killing vector fields on gradient Ricci solitons.
method Analyzes the structure of gradient Ricci solitons to estimate dimensions of Killing vector fields.
result Maximal dimension of Killing vector fields on irreducible non-trivial gradient Ricci solitons.
FPGs use structure to improve policy learning in complex tasks.
problem Policy gradient methods struggle with high-dimensional action spaces and objective multiplicity.
method Factor baseline and action-target influence network to reduce gradient variance.
result FPGs provide a general framework for state-of-the-art algorithms and improve performance.
It is inevitable to train large deep learning models on a large-scale cluster equipped with accelerators system. Deep gradient compression would highly increase the bandwidth utilization and speed up the training process but hard to implement on ring structure. In this paper, we find that redundant gradient and gradien…