Entropy rigidity for Finsler flows but collapse for Reeb flows.
problem Entropy behavior of Reeb and Finsler flows on contact manifolds.
method Analysis of topological entropy for Reeb and Finsler flows.
result Uniform positive lower bound for Finsler flows but arbitrarily small topological entropy for Reeb flows.
Study minimal volume entropy for free-by-cyclic groups and 2D right-angled Artin groups.
problem Characterize minimal volume entropy for aspherical simplicial complexes with these groups as fundamental groups.
method Algebraic and geometric characterization, using fiber π1-growth collapse and non-collapsing assumptions. result Provide bounds and criteria for minimal volume entropy in aspherical simplicial complexes.
GAN+VER improves GANs by regularizing entropy to reduce mode collapse.
problem Mode collapse in GANs where the generator fails to capture all modes.
method Maximizing a variational lower bound on the entropy of generated samples.
result Significant improvement in evaluation metrics for real and generated samples.
Self-attention networks localize when eigenspectrum variance is small.
problem Self-attention mechanisms can lead to rank and entropy collapses, reducing model expressivity and trainability.
method Characterized attention localization using query-key eigenspectrum variance.
result Small eigenspectrum variance prevents both rank and entropy collapses, improving model performance.
In this paper it is proven that the volume entropy of a riemannian metric evolving by the Ricci flow, if does not collapse, nondecreases. Therefore, it provides a sufficient condition for a solution to collapse. Then, for the limit solutions of type I or III, the limit entropy is the limit of the entropy as t approac…
Uniform entropy bound for Ricci shrinkers with bounded curvature.
problem Bounding entropy for Ricci shrinkers with specific curvature constraints.
method Establishing uniform entropy bounds for simply connected Ricci shrinkers with a finite second homotopy group and uniform curvature bounds.
result Uniform entropy bound for simply connected Ricci shrinkers with a finite second homotopy group and uniform curvature bounds.
This paper extends neural collapse to imbalanced data under cross-entropy loss.
problem Analyzing neural collapse in deep networks with imbalanced data.
method Using the unconstrained feature model and cross-entropy loss, the paper studies neural collapse in imbalanced datasets.
result Feature vectors within the same class collapse to a single mean vector, but angles between them depend on sample size.
New method prevents entropy collapse in Transformer training, leading to more stable and robust models.
problem Training instability in Transformers, especially in attention layers.
method Spectral normalization with a learned scalar to prevent entropy collapse.
result Prevents entropy collapse, leading to more stable training.
We localize the entropy functionals of G. Perelman and generalize his no-local-collapsing theorem and pseudo-locality theorem. Our generalization is technically inspired by further development of Li-Yau estimate along the Ricci flow. It can be used to show the Gromov-Hausdorff convergence of the Kähler Ricci flow on ea…
The study shows almost maximal volume entropy rigidity for certain manifolds with integral Ricci curvature.
problem Volume entropy rigidity for manifolds with lower integral Ricci curvature bound.
method Analyzing manifolds with specific integral Ricci curvature bounds, diameter, and volume entropy.
result The universal cover of the manifold is close to a hyperbolic space form under certain conditions.
COME replaces entropy minimization to prevent model collapse.
problem Overconfidence in entropy minimization leads to model collapse.
method COME explicitly models uncertainty with a Dirichlet prior distribution.
result COME achieves state-of-the-art performance on various TTA settings.
Paper explains neural collapse in neural networks using a new model.
problem Understanding neural collapse in neural networks during training.
method Introducing the unconstrained layer-peeled model (ULPM) to prove gradient flow convergence to critical points of a minimum-norm separation problem.
result Proves that all critical points are strict saddle points except the global minimizers exhibiting neural collapse.
We derive the entropy formula for the linear heat equaiton on complete Riemannian manifolds with nonnegative Ricci curvature. As applications, we study the relation between the value of entropy and the volume of balls of various scales. The results are simpler version, without Ricci flow, of Perelman's recent results o…
We show that if a closed manifold M admits an F-structure (possibly of rank 0) then its minimal entropy vanishes. In particular, this is the case if M admits a non-trivial circle action. As a corollary we obtain that the simplicial volume of a colsed manifold admitting an F-structure is zero. We also show that if M adm…
Entrocraft addresses RL performance saturation in LLMs by customizing entropy curves.
problem Performance saturation in RL algorithms for LLMs.
method Entrocraft uses rejection sampling to bias advantage distributions for customized entropy schedules.
result Entrocraft significantly improves generalization, output diversity, and long-term training in 4B models.
This paper extends neural collapse to class-imbalanced datasets using an unconstrained ReLU feature model.
problem Understanding neural collapse in class-imbalanced datasets with cross-entropy loss.
method Generalized neural collapse to class-imbalanced settings using an unconstrained ReLU feature model.
result Class-means converge to orthogonal vectors with different lengths, and classifier weights align to these vectors.
We make use of F-structures and technology developed by Paternain - Petean to compute minimal entropy, minimal volume, and Yamabe invariant of symplectic 4-manifolds, as well as to study their collapse with sectional curvature bounded from below. À la Gompf, we show that these invariants vanish on symplecti…
New regularization method reduces support of empirical risk minimization solutions.
problem Regularization in empirical risk minimization with relative entropy.
method Introduces Type-II regularization, characterizes solutions, analyzes properties of relative entropy.
result Type-II regularization collapses solution support into reference measure's support.
Fine-tunes diffusion models to generate diverse samples with high genuine rewards.
problem Reward collapse in finetuning diffusion models.
method Entropy-regularized control against pretrained diffusion models.
result Efficient generation of diverse samples with high genuine rewards.
Study shows translators can have non-removable singularities at infinity but eventually converge to unique planes.
problem Understanding singularities and convergence of translators at infinity.
method Global analysis of quasilinear soliton equations, sharp non-standard elliptic decay estimates, and potential theory.
result Finite entropy, finite genus translators converge to uniquely determined planes at infinity.
The paper extends entropy formulas to super Ricci flows on metric measure spaces.
problem Entropy formulas for super Ricci flows on metric measure spaces.
method Extending Perelman's W-entropy and Shannon entropy power to super Ricci flows. result Equivalence between volume non-local collapsing property and lower boundedness of W-entropy on RCD(0,N) spaces. In this short note, we analyze geometric properties of orbit spaces of certain involutions in dimensions four, five, and six. We consider constructions of F-structures on manifolds of dimension at least four that allows us to study minimal entropy, minimal volume, collapse with bounded curvature, and sign o…
A new loss function HUG decouples and generalizes neural collapse.
problem Neural collapse limits in deep learning models.
method Hyperspherical uniformity gap (HUG) as a unified framework.
result HUG decouples and generalizes neural collapse, improving model flexibility and robustness.
Study of geometric actions on CAT(0) spaces and their limits.
problem Understanding limits of geometric actions on CAT(0) spaces.
method Analysis of convergence and splitting/collapsing phenomena in CAT(0) lattices and orbispaces.
result Proof of compactness theorem for CAT(0) homology orbifolds.
New insights into CE dynamics reveal how Hadamard initialization simplifies softmax.
problem Understanding the dynamics of cross-entropy training loss in deep learning.
method Analyzing a two-layer linear neural network with standard-basis vectors as inputs.
result Gradient flow on cross-entropy converges to neural collapse geometry, proving global convergence.
We show vanishing results about the infimum of the topological entropy of the geodesic flow of homogeneous smooth four manifolds. We prove that any closed oriented geometric four manifold has zero minimal entropy if and only if it has zero simplicial volume. We also show that if a four manifold M admits a geometric dec…
Unified theory explains two failure modes of deep transformers and provides initialisation guidelines.
problem Two failure modes (rank collapse and entropy collapse) of self-attention layers in deep transformers.
method Analytical theory of signal propagation through deep transformers, using the Random Energy Model analogy.
result Simple algorithm to compute trainability diagrams for correct initialisation hyper-parameters.
We analyze neural collapse in neural networks, showing that features collapse to vertices of a Simplex ETF.
problem Understanding and optimizing the features learned in the last layer of neural networks during training.
method Simplified unconstrained feature model, studying the global optimization landscape of cross-entropy loss with weight decay.
result The global minimizers of the loss are Simplex ETFs, and other critical points are strict saddles with negative curvature.
Proposes MEDM to balance entropy minimization and diversity maximization for better domain adaptation.
problem Trivial solutions in entropy minimization for unsupervised domain adaptation.
method Introduces diversity maximization to balance with entropy minimization, controlled by deep embedded validation.
result MEDM outperforms state-of-the-art methods on four domain adaptation datasets.
In 1870s, L. Boltzmann proved the famous H-theorem for the Boltzmann equation in the kinetic theory of gas and gave the statistical interpretation of the thermodynamic entropy. In 2002, G. Perelman introduced the notion of W-entropy and proved the W-entropy formula for the Ricci flow. This plays a crucial role in…
The paper extends Perelman's theorems on Ricci flow entropy.
problem Understanding the behavior of Ricci flow under various conditions.
method Localization of entropy functionals and development of Li-Yau estimates.
result Generalization of Perelman's no-local-collapsing and pseudo-locality theorems.
Generative adversarial networks (GANs) are a powerful approach to unsupervised learning. They have achieved state-of-the-art performance in the image domain. However, GANs are limited in two ways. They often learn distributions with low support---a phenomenon known as mode collapse---and they do not guarantee the exist…
Quantitative rigidity theorem for Alexandrov spaces with curvature bounds.
problem Quantifying rigidity in Alexandrov spaces with curvature constraints.
method Using Gromov-Hausdorff distance and properties of Alexandrov spaces.
result Alexandrov spaces with curvature bounds are close to hyperbolic manifolds.
In this note, we prove a uniform distance distortion estimate for Ricci flows with uniformly bounded scalar curvature, independent of the lower bound of the initial μ-entropy. Our basic principle tells that once correctly renormalized, the metric-measure quantities obey similar estimates as in the non-collapsing case…
We study the problem of existence of F-structures on compact complex surfaces, giving a complete classification modulo the gap in the classification of surfaces of class VII. We then use these results to study the minimal entropy problem for compact complex surfaces. For instance we prove that compact Kahler surfaces o…
Our research proves neural collapse in deep ResNets and transformers is globally optimal.
problem Understanding neural collapse in deep learning models.
method Analysis of deep regularized transformers and ResNets trained with cross entropy or mean squared error loss.
result Global optima of deep regularized transformers and ResNets are approximately collapsed, becoming more prominent as depth increases.
Study shows only grim reaper cylinder for certain self-translating surfaces.
problem Characterizing self-translating surfaces in 3D space.
method Used parabolicity in a weighted setting and universally L-superharmonic functions.
result Characterized the grim reaper cylinder as the only finite entropy self-translating 2-surface in R^3 of width π and bounded from below.
Entropy asymmetry affects regularization in ERM, leading to biased solutions.
problem Analyzing the impact of relative entropy asymmetry in ERM regularization.
method Examined Type-I and Type-II ERM-RER, comparing their solutions and properties.
result Type-II ERM-RER regularization introduces a strong bias against training data.
MIND estimates mutual information from ordinal data without full distributional knowledge.
problem Estimating mutual information from ordinal data with limited data.
method Copula-based maximum-entropy estimation of copula entropies.
result MIND estimator is consistent, data-efficient, and unbounded for any sample size.
AER dynamically adjusts entropy regularization for better LLM reinforcement learning.
problem Policy entropy collapse in RLVR training limits exploration and reasoning performance.
method Adaptive Entropy Regularization (AER) with difficulty-aware coefficient allocation, initial-anchored target entropy, and dynamic global coefficient adjustment.
result AER consistently outperforms baselines on mathematical reasoning benchmarks, improving both accuracy and exploration.
This work justifies neural collapse under MSE loss and analyzes the optimization landscape.
problem Understanding neural collapse in deep neural networks under MSE loss.
method Global landscape analysis of vanilla nonconvex MSE loss.
result The only global minimizers are neural collapse solutions.
HCLM framework uses entropy regularization for open learning systems.
problem Real-world AI challenges and limitations of deep learning.
method Dynamical and information-theoretic framework with entropy regularization.
result Geometric entropy surrogates, especially log-determinant covariance entropy, induce stronger and more stable information forces.
Deep nets trained with MSE loss exhibit Neural Collapse, collapsing features and classifiers to class means.
problem Understanding Neural Collapse in MSE-trained deep nets.
method Developed a new MSE loss decomposition and introduced the central path concept.
result Exact dynamics of Neural Collapse along the central path can be predicted.
Study shows neural collapse is invariant to class imbalances under certain conditions.
problem Neural collapse properties are only valid for balanced data.
method Adopted UFM and introduced SELI for invariant characterization.
result Embeddings and classifiers always interpolate a simplex-encoded label matrix regardless of class imbalances.
Persistent entropy detects phase transitions in complex systems.
problem Detecting phase transitions in complex systems.
method Established a general theorem for persistent entropy to reliably detect phase transitions, introduced operational framework for finite-time computations.
result Persistent entropy exhibits an asymptotically non-vanishing gap across phases, robust numerical signatures across experiments.
Deep linear networks exhibit collapsing features and classifiers across datasets.
problem Understanding the collapse of features and classifiers in deep linear networks.
method Theoretical and empirical analysis of deep linear networks with MSE and CE losses.
result Deep linear networks exhibit NC properties, collapsing features and classifiers to orthogonal vectors.
New method prevents mode collapse in deep SVDD for anomaly detection.
problem Mode collapse in deep SVDD due to architectural constraints.
method Two regularizers: noise injection and minibatch variance penalization.
result Regularized deep SVDD outperforms state-of-the-art methods.
AdaDEM decouples EM into two parts to improve class overlap and uncertainty.
problem Improper EM limits its effectiveness in various machine learning tasks.
method Decouple EM into CADF and GMC, and AdaDEM normalizes CADF reward and uses MEC.
result AdaDEM outperforms classical EM and improves performance in noisy and dynamic environments.