Entropy norm on surface diffeomorphisms is unbounded.
problem Bounding entropy norm on surface diffeomorphisms.
method Analyzing the group of diffeomorphisms of a closed orientable surface.
result Entropy norm is unbounded on the group of diffeomorphisms of a closed orientable surface.
Innovates volume entropy semi-norm, proving equivalence to simplicial volume.
problem Equivalence of volume entropy and simplicial volume in real homology.
method Introduces volume entropy semi-norm and proves its equivalence to simplicial volume.
result Equivalence of volume entropy semi-norm and simplicial volume in every dimension.
Proposes a new criterion for selecting Nash equilibria considering both utility and inequality.
problem Finding a fair Nash equilibrium in group decision-making.
method Introduces entropy-norm space for geometric selection of strict Nash equilibria.
result The closest entropy-norm pair to the largest entropy-norm pair in rescaled space is the most suitable equilibrium.
Note proves bi-Lipschitz embeddings for Ham(S^2) metrics.
problem Entropy norms on Ham(S^2) bi-Lipschitz embeddings.
method Proves bi-Lipschitz embeddings for Ham(S^2) with entropy metrics.
result Existence of bi-Lipschitz embeddings for Ham(S^2) metrics.
Study on automorphisms of K3 and Enriques surfaces, proving entropy gaps and achirality.
problem Entropy norms and achirality of automorphisms on K3 and Enriques surfaces.
method Proves gap theorems for entropy norms and studies achirality in terms of genus-one fibrations.
result Entropy gaps and achirality results for automorphisms of K3 and Enriques surfaces.
Paper connects loss functions and t-norms for faster deep learning convergence.
problem Improving convergence rates in deep learning models.
method Interprets loss functions through t-norms and generator functions.
result Derives a general relation between loss functions and t-norms leading to faster convergence.
The paper proves inequalities linking geometric norms and Thurston norms in hyperbolic 3-manifolds.
problem Inequalities linking geometric norms and Thurston norms in hyperbolic 3-manifolds.
method Analyzes geometric L2-norms, Thurston norms, and Lipschitz maps to prove inequalities. result Proves an inequality between geometric L2-norm and Thurston norm, qualitatively sharp. This paper studies neural networks with bounded norms to avoid the curse of dimensionality.
problem The curse of dimensionality in approximating functions by neural networks.
method Investigates over-parameterized two-layer neural networks with norm constraints in RKHS.
result Improved sample complexity and generalization bounds for neural networks with bounded norms.
Thanks to a theorem of Brock on comparison of Weil-Petersson translation distances and hyperbolic volumes of mapping tori for pseudo-Anosovs, we prove that the entropy of a surface automorphism in general has linear bounds in terms of Gromov norm of its mapping torus from below and in bounded geometry case from above. …
Characterizes sample complexity for outcome indistinguishability in machine learning.
problem Outcome indistinguishability in machine learning, focusing on distinguishers and predictors.
method Sample complexity characterized by metric entropy of predictor and distinguisher classes, using dual Minkowski norms.
result Equivalence and tightness of sample complexity characterizations in distribution-specific and distribution-free settings.
Simple regional perturbations maintain model transferability while reducing adversarial example distortion.
problem Comparing efficacy of regional adversarial attacks without complex methods.
method Developed a simple regional adversarial perturbation attack using cross-entropy sign.
result Localized adversarial examples require significantly less Lp norm distortion compared to non-local counterparts. We consider a hyperbolic surface bundle over the circle with the smallest known volume among hyperbolic manifolds having 3 cusps, so called "the magic manifold". We compute the entropy function on the fiber face of the unit ball with respect to the Thurston norm, determine homology classes whose representatives are gen…
Study simplicial volume for fixed fundamental groups, finding gaps.
problem Understanding simplicial volume for manifolds with fixed fundamental group.
method Relate gap problem to rationality questions in bounded (co)homology.
result Show existence of gaps in simplicial volume spectrum at zero.
Sharp lower bounds on shallow neural networks' approximation rates are derived.
problem The efficiency of shallow neural networks in approximating functions.
method Lower bounding the L2-metric entropy and Kolmogorov n-widths of the convex hull of neural network basis functions. result Sharp lower bounds on the approximation rates for shallow neural networks are provided.
Algorithm computes Thurston norm for hyperbolic 3-manifolds.
problem Computing the Thurston norm for hyperbolic 3-manifolds.
method Developed a theory of spun-normal immersed surfaces and implemented an algorithm.
result Computed the unit ball of the Thurston norm for cusped hyperbolic 3-manifolds.
Quantum ML predicts data with improved speed and accuracy.
problem Predicting data using maximum likelihood in a quantum setting.
method Quantum states embedding and minimization of quantum relative entropy.
result Unified framework for classical and quantum LLMs with performance guarantees.
Paper improves Frank-Wolfe algorithm's efficiency bounds.
problem Establishing efficient iteration complexity for Frank-Wolfe algorithm.
method Using metric entropy to provide lower bounds.
result Frank-Wolfe requires many iterations for certain problems.
This paper provides a general result on controlling local Rademacher complexities, which captures in an elegant form to relate the complexities with constraint on the expected norm to the corresponding ones with constraint on the empirical norm. This result is convenient to apply in real applications and could yield re…
We consider the minimum error entropy (MEE) criterion and an empirical risk minimization learning algorithm in a regression setting. A learning theory approach is presented for this MEE algorithm and explicit error bounds are provided in terms of the approximation ability and capacity of the involved hypothesis space w…
The relaxed maximum entropy problem is concerned with finding a probability distribution on a finite set that minimizes the relative entropy to a given prior distribution, while satisfying relaxed max-norm constraints with respect to a third observed multinomial distribution. We study the entire relaxation path for thi…
New microcanonical models approximate non-Gaussian processes with long-range correlations.
problem Approximating non-Gaussian stationary processes with long-range correlations.
method Microcanonical models conditioned by energy vector, gradient descent, multiscale energy vectors.
result Microcanonical gradient descent processes converge and capture sparsity.
Paper explains neural collapse in neural networks using a new model.
problem Understanding neural collapse in neural networks during training.
method Introducing the unconstrained layer-peeled model (ULPM) to prove gradient flow convergence to critical points of a minimum-norm separation problem.
result Proves that all critical points are strict saddle points except the global minimizers exhibiting neural collapse.
New method prevents entropy collapse in Transformer training, leading to more stable and robust models.
problem Training instability in Transformers, especially in attention layers.
method Spectral normalization with a learned scalar to prevent entropy collapse.
result Prevents entropy collapse, leading to more stable training.
In this paper, we consider low rank matrix estimation using either matrix-version Dantzig Selector A^λd or matrix-version LASSO estimator A^λL. We consider sub-Gaussian measurements, i.e., the measurements X1,…,Xn∈Rm×m have i.i.d. sub-Gaussian entries. Suppose $\textrm…
Muon replaces matrix gradient with polar factor, optimizing flat spectrum updates
problem Optimization bias in matrix updates
method Using polar factor of gradient
result Muon update maximizes entropy among bounded updates
Lower bound shows no acceleration for specific convex optimization class.
problem Proving lower bounds for convergence rates of convex optimization methods.
method Proving Ω(L/T) lower bound for minimization of convex and L-smooth functions relative to negative entropy. result Mirror descent is optimal up to a logarithmic factor in the class of functions considered.
Proves uniform K-stability is open in Kähler cone.
problem Stability of Kähler metrics in complex geometry.
method Introduced new norm on test configurations and estimates for non-archimedean energy functionals.
result Uniform K-stability is an open condition in the Kähler cone.
We consider the problem of online nonparametric regression with arbitrary deterministic sequences. Using ideas from the chaining technique, we design an algorithm that achieves a Dudley-type regret bound similar to the one obtained in a non-constructive fashion by Rakhlin and Sridharan (2014). Our regret bound is expre…
New algorithms improve neural network generalization by finding flat minima.
problem Finding better generalization in neural networks through flat minima.
method Developed Entropy-SGD and Replicated-SGD algorithms to maximize flatness in the loss function.
result Consistently improved generalization error for various deep learning architectures.
New scalable algorithm for non-negative linear regression with entropy-regularized OT loss.
problem Generalizing task-specific linear models to broader applications.
method Sinkhorn-like scaling iterations for convex penalty and datafit terms.
result Simple multiplicative updates for various penalty and datafit terms.
Mirror descent algorithm recovers low-rank matrices in matrix sensing.
problem Matrix sensing with low-rank matrices under certain conditions.
method Discrete-time mirror descent applied to empirical risk with Bregman divergence analysis.
result Mirror descent converges to a matrix minimizing a specific nuclear norm-related quantity.
This paper proves a generalization bound for complex-valued neural networks scaling with spectral complexity.
problem Ensuring the performance of complex-valued neural networks on unseen data.
method Theoretical derivation using Maurey Sparsification Lemma and Dudley Entropy Integral, empirical validation on various datasets.
result The spectral complexity of weight matrices is a significant factor in the generalization ability of complex-valued neural networks.
The density matrices are positively semi-definite Hermitian matrices of unit trace that describe the state of a quantum system. The goal of the paper is to develop minimax lower bounds on error rates of estimation of low rank density matrices in trace regression models used in quantum state tomography (in particular, i…
Gradient descent biases linear models in next-token prediction towards data entropy.
problem Optimization bias in next-token prediction models.
method Analysis of gradient descent on linear models with sparse conditional distributions.
result Gradient descent selects parameters that equate token logits differences to log-odds in the data subspace.
Recent contributions have framed linear system identification as a nonparametric regularized inverse problem. Relying on ℓ2-type regularization which accounts for the stability and smoothness of the impulse response to be estimated, these approaches have been shown to be competitive w.r.t classical parametric met…
Study bounds on kernel function entropy for finite measures.
problem Investigate bounds on the ε-entropy of kernel classes.
method Sharp upper and lower bounds for p in [1, +∞] derived from eigenvalue behavior and Mercer series convergence.
result Proves tighter bounds for general kernels compared to previous work.
A method for learning from unlabeled time-series data using temporal smoothing and entropy maximization.
problem Learning from unlabeled time-series data efficiently and accurately.
method Training a feedforward neural network with two objectives: temporal smoothing and entropy maximization.
result The method extracts slowly evolving information from time-series data, filtering out noise.
We study the problem of reconstructing an unknown matrix M of rank r and dimension d using O(rd poly log d) Pauli measurements. This has applications in quantum state tomography, and is a non-commutative analogue of a well-known problem in compressed sensing: recovering a sparse vector from a few of its Fourier coeffic…
Study on convergence of exponential probability measures with applications to maximum entropy models and SGLD.
problem Characterizing the limit of probability measures with exponential densities as temperature approaches zero.
method Quantitative bounds on Wasserstein distance using geometric measure theory tools.
result Established quantitative convergence results for norm-like potentials under invertibility conditions.
The paper analyzes how well classes are separated in neural network feature space.
problem Understanding class separability in neural network feature space.
method Theoretical analysis of intra-class and inter-class distances in feature space.
result A lower bound for the probability of inter-class distance being greater than intra-class distance as a function of loss value.
We present a method for training multi-label, massively multi-class image classification models, that is faster and more accurate than supervision via a sigmoid cross-entropy loss (logistic regression). Our method consists in embedding high-dimensional sparse labels onto a lower-dimensional dense sphere of unit-normed …
Study shows fast rates for inverse reinforcement learning with linear rewards.
problem Entropy-regularized min-max inverse reinforcement learning in finite-horizon MDPs.
method Structural and statistical analysis of Min-Max-IRL with pseudo-self-concordance.
result Both trajectory-level KL divergence and parameter error decay at O(n−1). The paper explores how neural networks avoid overfitting by learning from flat minima.
problem How to prevent overfitting in neural networks by learning from flat minima.
method Study of one- and two-layer neural network models, derivation of algorithms focusing on wide flat regions.
result Wide flat minima coexist with narrower minima and critical points, and are associated with good minimizers.
Cognitive abstraction metrics predict deep neural networks' generalization ability.
problem Understanding deep neural networks' generalization ability.
method Defined Cognitive Neural Activation (CNA) metric based on information complexity and activation patterns.
result CNA is highly predictive of generalization ability, outperforming other metrics.
A closed four dimensional manifold cannot possess a non-flat Ricci soliton metric with arbitrarily small L2-norm of the curvature. In this paper, we localize this fact in the case of shrinking Ricci solitons by proving an ε-regularity theorem, thus confirming a conjecture of Cheeger-Tian. As applications…
New research shows SAA can outperform SA for Wasserstein barycenters.
problem Optimizing Wasserstein barycenters with entropy regularization.
method Comparison of Stochastic Approximation (SA) and Sample Average Approximation (SAA) for large-scale problems.
result SAA can be more efficient than SA for Wasserstein barycenters, especially in large-scale settings.
Logit dynamics formula reveals self-regulation in softmax policy gradient methods.
problem Understanding the stability and convergence of softmax policy gradient methods.
method Deriving the exact formula for the L2 norm of the logit update vector.
result Logit update magnitudes are modulated by action probability and policy concentration.
A new method speeds up SoftMax normalization for embedding learning.
problem Efficiently learning distributed representations with SoftMax normalization.
method Proposes a linear-time heuristic approximation for mSoftMax(XYT), optimizing cross entropy. result Achieves higher or comparable accuracy to existing methods with lower computational time.