AR-DAE approximates entropy gradient for machine learning models.
problem Intractable computation of entropy gradient for continuous distributions.
method Amortized residual denoising autoencoder (AR-DAE) to approximate entropy gradient.
result AR-DAE provides an unbiased gradient approximation for entropy.
Entropy regularization improves policy optimization in reinforcement learning.
problem Improving policy optimization in reinforcement learning.
method Entropy regularization is introduced to soften the greedy policy towards a more diverse softmax policy, leading to a continuously parameterized algorithm that interpolates between policy gradient and Q-learning.
result An intermediate algorithm can improve performance in reinforcement learning.
We study a Boltzmann's type entropy functional (which appeared in existing literature) defined on Kähler metrics of a fixed Kähler class. The critical points of this functional are gradient Kähler-Ricci solitons, and the functional was known to be monotonically increasing along the Kähler-Ricci flow in the canonical cl…
The paper establishes sub-gradient estimates and entropy formulas for quaternionic contact geometry heat equations.
problem Developing sub-gradient estimates and entropy formulas for quaternionic contact geometry.
method Establishing sub-gradient estimates and entropy formulas for the quaternionic contact heat equation.
result Two Perelman-type entropy formulas and sub-gradient estimates for the quaternionic contact heat equation.
Abstract: Necessary and sufficient conditions for gradient flows of relative entropy in Lindblad equations.
problem Conditions for gradient flows in finite-dimensional Lindblad equations.
method Analyzes conditions for a finite-dimensional Lindblad equation to have a gradient flow structure for the von Neumann relative entropy.
result A finite-dimensional Lindblad equation admits a gradient flow structure for the von Neumann relative entropy if and only if the BKM-detailed balance condition holds.
Entropy-regularized NPG converges linearly with linear function approximation.
problem Analyzing convergence of entropy-regularized NPG with function approximation.
method Established finite-time convergence analyses with entropy regularization and linear function approximation.
result Entropy-regularized NPG achieves linear convergence up to a function approximation error.
Gradient flow expands curves to round shapes.
problem Expanding curves to round shapes.
method Steepest descent L2-gradient flow of entropy.
result Flow converges to a round expanding circle for various initial curves.
Study shows policy gradient convergence for entropy-regularized MDPs with neural nets in mean-field regime.
problem Global convergence of policy gradient for entropy-regularized MDPs with neural network approximation.
method Softmax policy with neural network approximation in mean-field regime, gradient flow in 2-Wasserstein metric, exponential convergence under sufficient regularization.
result Gradient flow converges exponentially fast to the unique stationary solution under sufficient regularization.
A new method for sampling high-dimensional distributions overcomes overfitting.
problem Overfitting in energy-based models during gradient descent.
method Mean-field microcanonical gradient descent, which samples multiple data points simultaneously.
result The method reduces entropy loss while maintaining likelihood fit, improving overfitting issues.
Structured entropy improves classification performance on structured targets.
problem Cross-entropy loss fails to account for target variable structure.
method Proposes structured entropy, a generalization of entropy using random partitions.
result Structured cross-entropy loss yields better results on classification problems with known structure.
Unified framework for analyzing gradient flows of measures with exponential decay of entropy.
problem Analyzing exponential decay of entropy functionals in gradient flows of measures.
method Characterization of global exponential decay behaviors using Hellinger-Kantorovich geometry, shape-mass decomposition, and Polyak-Łojasiewicz-type inequalities.
result Unified theoretical framework for gradient flows with complete analysis of exponential decay behaviors.
Estimate relaxation times in nonextensive systems using gradient flow for Tsallis entropy maximization.
problem Estimating relaxation times in financial market dynamics.
method Developing a method using EGF for maximizing Tsallis entropy.
result Longer relaxation times for nonextensive systems compared to Shannon entropy.
Entropy regularization improves MFG learning efficiency and stability.
problem Improving Mean Field Game learning efficiency and stability.
method Entropy regularization applied to MFG with learning.
result Entropy regularization yields time-dependent policies and stabilizes convergence.
In recent years, deep reinforcement learning has been shown to be adept at solving sequential decision processes with high-dimensional state spaces such as in the Atari games. Many reinforcement learning problems, however, involve high-dimensional discrete action spaces as well as high-dimensional state spaces. This pa…
Gradient flow in softmax models tends to produce low-entropy outputs.
problem Understanding the training dynamics of softmax-based models.
method Analysis of gradient flow dynamics in the value-softmax model.
result Gradient flow drives optimization towards low-entropy solutions.
Proposes MEDM to balance entropy minimization and diversity maximization for better domain adaptation.
problem Trivial solutions in entropy minimization for unsupervised domain adaptation.
method Introduces diversity maximization to balance with entropy minimization, controlled by deep embedded validation.
result MEDM outperforms state-of-the-art methods on four domain adaptation datasets.
Softmax policy gradient methods converge at O(1/t) rate with constants depending on problem and initialization.
problem Understanding convergence rates of softmax policy gradient methods in tabular settings.
method Analysis of softmax policy gradient and entropy regularized policy gradient methods, using Łojasiewicz inequality and lower bounds.
result Entropy regularization improves convergence rate from O(1/t) to O(e−c⋅t). A new measure helps compute suboptimality in entropy-regularized methods.
problem Computing suboptimality in entropy-regularized variational objectives when unnormalised densities are unavailable.
method Introduced 'kernel gradient discrepancy' (KGD) to compute suboptimality explicitly.
result KGD characterizes kernel Stein discrepancy (KSD) in the standard Bayesian context and measures variational gradient size.
REGS samples from unnormalized distributions using gradient flow and neural networks.
problem Sampling from unnormalized distributions with high accuracy and efficiency.
method REGS is a particle method that iteratively transforms samples from a reference distribution to match an unnormalized target distribution using Wasserstein gradient flow and neural networks.
result REGS outperforms state-of-the-art methods in sampling from challenging multimodal distributions and real datasets.
Entropy-regularized NPG methods converge linearly in discounted MDPs.
problem Theoretical limitations of NPG methods in reinforcement learning.
method Entropy regularization in conjunction with NPG methods for discounted MDPs.
result Entropy-regularized NPG methods converge linearly in discounted MDPs.
Paper proves volume growth estimate for steady gradient Ricci solitons.
problem Estimating the volume growth of steady gradient Ricci solitons.
method Proved a volume growth estimate using Nash entropy.
result Volume growth rate is no smaller than $r^{rac{n+1}{2}}$.
Maximum entropy deep reinforcement learning (RL) methods have been demonstrated on a range of challenging continuous tasks. However, existing methods either suffer from severe instability when training on large off-policy data or cannot scale to tasks with very high state and action dimensionality such as 3D humanoid l…
ESPO optimizes LLMs for complex tasks by balancing fine-grained updates and stability.
problem Gradient underutilization in sequence-level optimization.
method Entropy Grouping Importance Sampling and Entropy Adaptive Clipping.
result ESPO accelerates convergence and achieves state-of-the-art performance.
Trust-region methods have yielded state-of-the-art results in policy search. A common approach is to use KL-divergence to bound the region of trust resulting in a natural gradient policy update. We show that the natural gradient and trust region optimization are equivalent if we use the natural parameterization of a st…
HCLM framework uses entropy regularization for open learning systems.
problem Real-world AI challenges and limitations of deep learning.
method Dynamical and information-theoretic framework with entropy regularization.
result Geometric entropy surrogates, especially log-determinant covariance entropy, induce stronger and more stable information forces.
In his 2011 work, Maas has shown that the law of any time-reversible continuous-time Markov chain with finite state space evolves like a gradient flow of the relative entropy with respect to its stationary distribution. In this work we show the converse to the above by showing that if the relative law of a Markov chain…
The purpose of this work is to study some monotone functionals of the heat kernel on a complete Riemannian manifold with nonnegative Ricci curvature. In particular, we show that on these manifolds, the gradient estimate of Li and Yau, the gradient estimate of Ni, the monotonicity of the Perelman's entropy and the volum…
We consider gradient estimates to positive solutions of porous medium equations and fast diffusion equations: ut=Δφ(up) associated with the Witten Laplacian on Riemannian manifolds. Under the assumption that the m-dimensional Bakry-Emery Ricci curvature is bounded from below, we obtain gradient estimates which…
Proves convergence of gradient Ricci shrinkers with uniform bounds.
problem Compactness and energy concentration in gradient Ricci shrinkers.
method Bubble-tree convergence and local energy analysis.
result No energy concentrates in neck regions, leading to a local diffeomorphism finiteness theorem.
Entropy measures geodesic flow complexity.
problem Measuring complexity of geodesic flows on manifolds.
method Introduced barcode entropy to measure exponential growth rate of not-too-short bars in Morse-theoretic barcodes.
result Barcode entropy bounds topological entropy and vice versa.
New method accelerates convergence for entropy-regularized reinforcement learning problems.
problem Slow convergence of standard first-order methods for entropy-regularized Markov decision processes.
method Introduce a quadratically convexified primal-dual formulation and a new interpolating metric to accelerate convergence.
result Global convergence and exponential convergence rate for the new method.
In this paper we discuss Perelman's Lambda-functional, Perelman's Ricci shrinker entropy as well as the Ricci expander entropy on a class of manifolds with isolated conical singularities. On such manifolds, a singular Ricci de Turck flow preserving the isolated conical singularities exists by our previous work. We prov…
We study blow-ups around fixed points at Type I singularities of the Ricci flow on closed manifolds using Perelman's W-functional. First, we give an alternative proof of the result obtained by Naber and Enders-Müller-Topping that blow-up limits are non-flat gradient shrinking Ricci solitons. Our second and main result …
Softmax policy gradient achieves global optimality in wide neural networks with entropy regularization.
problem Optimizing softmax policies with neural networks in the mean-field regime.
method Modeling neural networks as Wasserstein gradient flows and proving global optimality of fixed points.
result Global optimality of softmax policy gradient in wide single hidden layer neural networks with entropy regularization.
The paper tackles scalarization issues in A2C RL algorithms, proposing methods to avoid gradient overlap and noise.
problem Scalarization issues in A2C RL algorithms leading to gradient overlap and uncontrolled noise.
method Proposes techniques to avoid gradient overlap and noise in A2C RL algorithms.
result Pilot experiments show the proposed method speeds up training in A2C RL algorithms.
We introduce Implicit Policy, a general class of expressive policies that can flexibly represent complex action distributions in reinforcement learning, with efficient algorithms to compute entropy regularized policy gradients. We empirically show that, despite its simplicity in implementation, entropy regularization c…
In this paper, we propose an implicit gradient descent algorithm for the classic k-means problem. The implicit gradient step or backward Euler is solved via stochastic fixed-point iteration, in which we randomly sample a mini-batch gradient in every iteration. It is the average of the fixed-point trajectory that is c…
New analysis shows how cross-entropy training shapes attention in transformers.
problem Understanding how gradient-based learning creates the required internal geometry in transformers.
method Developed a first-order analysis of cross-entropy training effects on attention scores and values in a transformer attention head.
result Introduced an advantage-based routing law and responsibility-weighted update for attention scores and values, respectively.
Mathematical analysis of SNE and t-SNE for dimension reduction.
problem Optimal mapping of high-dimensional data to low dimensions.
method Gradient flow of relative entropy to minimize the distance between points.
result The diameter of the evolving sets remains bounded for SNE but may blow up for t-SNE.
The policy gradient theorem is defined based on an objective with respect to the initial distribution over states. In the discounted case, this results in policies that are optimal for one distribution over initial states, but may not be uniformly optimal for others, no matter where the agent starts from. Furthermore, …
The paper extends Ricci flow theory with Type-I scalar curvature bounds, proving entropy convergence and characterizing singular sets.
problem Extending Ricci flow theory with Type-I scalar curvature bounds.
method Type-I rescaling procedure and entropy analysis of conjugate heat kernels.
result Entropy of Ricci flow solutions converges to soliton entropy, characterizing singular sets.
The aim of this paper is to provide new theoretical and computational understanding on two loss regularizations employed in deep learning, known as local entropy and heat regularization. For both regularized losses we introduce variational characterizations that naturally suggest a two-step scheme for their optimizatio…
This paper proposes a new optimization algorithm called Entropy-SGD for training deep neural networks that is motivated by the local geometry of the energy landscape. Local extrema with low generalization error have a large proportion of almost-zero eigenvalues in the Hessian with very few positive or negative eigenval…
This paper improves Bayesian inference for predictive models with limited data.
problem Effective uncertainty quantification for training predictive models with limited data.
method Entropy-regularized gradient estimators to approximate the Bayesian posterior.
result The method generates diverse samples from the posterior distribution efficiently.
Unique continuation result for expanding Ricci solitons.
problem Unique continuation of expanding Ricci solitons.
method Optimal relative integral convergence rate, relative entropy.
result Well-defined relative entropy for expanding solitons.
Gradient descent biases linear models in next-token prediction towards data entropy.
problem Optimization bias in next-token prediction models.
method Analysis of gradient descent on linear models with sparse conditional distributions.
result Gradient descent selects parameters that equate token logits differences to log-odds in the data subspace.
New method trains normalizing flows using entropy-regularized transport.
problem Training continuous normalizing flows efficiently.
method Formulates flows as gradients of scalar potentials, training only these potentials.
result Trains normalizing flows without explicit flow computation during training.
Let K be an irreducible and reversible Markov kernel on a finite set X. We construct a metric W on the set of probability measures on X and show that with respect to this metric, the law of the continuous time Markov chain evolves as the gradient flow of the entropy. This result is a discrete counterpart of the Wassers…