We introduce the volume entropy semi-norm in real homology and show that it satisfies functorial properties similar to the ones of the simplicial volume. Answering a question of M. Gromov, we prove that the volume entropy semi-norm is equivalent to the simplicial volume semi-norm in every dimension. We also establish a…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes a new criterion for selecting Nash equilibria considering both utility and inequality.
We prove that the entropy norm on the group of diffeomorphisms of a closed orientable surface of positive genus is unbounded.
Study on automorphisms of K3 and Enriques surfaces, proving entropy gaps and achirality.
The paper proves inequalities linking geometric norms and Thurston norms in hyperbolic 3-manifolds.
In this note we prove that for each positive integer there exists a bi-Lipschitz embedding , where is equipped with the entropy metric. In particular, the same result holds when the entropy metric is substituted with the autonomous metric.
Deep learning has been shown to achieve impressive results in several domains like computer vision and natural language processing. A key element of this success has been the development of new loss functions, like the popular cross-entropy loss, which has been shown to provide faster convergence and to reduce the vani…
This paper studies neural networks with bounded norms to avoid the curse of dimensionality.
Thanks to a theorem of Brock on comparison of Weil-Petersson translation distances and hyperbolic volumes of mapping tori for pseudo-Anosovs, we prove that the entropy of a surface automorphism in general has linear bounds in terms of Gromov norm of its mapping torus from below and in bounded geometry case from above. …
Characterizes sample complexity for outcome indistinguishability in machine learning.
We consider a hyperbolic surface bundle over the circle with the smallest known volume among hyperbolic manifolds having 3 cusps, so called "the magic manifold". We compute the entropy function on the fiber face of the unit ball with respect to the Thurston norm, determine homology classes whose representatives are gen…
Simple regional perturbations maintain model transferability while reducing adversarial example distortion.
Study simplicial volume for fixed fundamental groups, finding gaps.
Sharp lower bounds on shallow neural networks' approximation rates are derived.
Algorithm computes Thurston norm for hyperbolic 3-manifolds.
Quantum ML predicts data with improved speed and accuracy.
Paper improves Frank-Wolfe algorithm's efficiency bounds.
This paper provides a general result on controlling local Rademacher complexities, which captures in an elegant form to relate the complexities with constraint on the expected norm to the corresponding ones with constraint on the empirical norm. This result is convenient to apply in real applications and could yield re…
We consider the minimum error entropy (MEE) criterion and an empirical risk minimization learning algorithm in a regression setting. A learning theory approach is presented for this MEE algorithm and explicit error bounds are provided in terms of the approximation ability and capacity of the involved hypothesis space w…
We study approximations of non-Gaussian stationary processes having long range correlations with microcanonical models. These models are conditioned by the empirical value of an energy vector, evaluated on a single realization. Asymptotic properties of maximum entropy microcanonical and macrocanonical processes and the…
The relaxed maximum entropy problem is concerned with finding a probability distribution on a finite set that minimizes the relative entropy to a given prior distribution, while satisfying relaxed max-norm constraints with respect to a third observed multinomial distribution. We study the entire relaxation path for thi…
New method prevents entropy collapse in Transformer training, leading to more stable and robust models.
Paper explains neural collapse in neural networks using a new model.
In this paper, we consider low rank matrix estimation using either matrix-version Dantzig Selector or matrix-version LASSO estimator . We consider sub-Gaussian measurements, , the measurements have sub-Gaussian entries. Suppose $\textrm…
Muon replaces matrix gradient with polar factor, optimizing flat spectrum updates
Proves uniform K-stability is open in Kähler cone.
We consider the problem of online nonparametric regression with arbitrary deterministic sequences. Using ideas from the chaining technique, we design an algorithm that achieves a Dudley-type regret bound similar to the one obtained in a non-constructive fashion by Rakhlin and Sridharan (2014). Our regret bound is expre…
New algorithms improve neural network generalization by finding flat minima.
New scalable algorithm for non-negative linear regression with entropy-regularized OT loss.
Mirror descent algorithm recovers low-rank matrices in matrix sensing.
This paper proves a generalization bound for complex-valued neural networks scaling with spectral complexity.
The density matrices are positively semi-definite Hermitian matrices of unit trace that describe the state of a quantum system. The goal of the paper is to develop minimax lower bounds on error rates of estimation of low rank density matrices in trace regression models used in quantum state tomography (in particular, i…
Gradient descent biases linear models in next-token prediction towards data entropy.
Recent contributions have framed linear system identification as a nonparametric regularized inverse problem. Relying on -type regularization which accounts for the stability and smoothness of the impulse response to be estimated, these approaches have been shown to be competitive w.r.t classical parametric met…
Study bounds on kernel function entropy for finite measures.
We study the problem of reconstructing an unknown matrix M of rank r and dimension d using O(rd poly log d) Pauli measurements. This has applications in quantum state tomography, and is a non-commutative analogue of a well-known problem in compressed sensing: recovering a sparse vector from a few of its Fourier coeffic…
Study on convergence of exponential probability measures with applications to maximum entropy models and SGLD.
In this paper, we focus on the separability of classes with the cross-entropy loss function for classification problems by theoretically analyzing the intra-class distance and inter-class distance (i.e. the distance between any two points belonging to the same class and different classes, respectively) in the feature s…
We present a method for training multi-label, massively multi-class image classification models, that is faster and more accurate than supervision via a sigmoid cross-entropy loss (logistic regression). Our method consists in embedding high-dimensional sparse labels onto a lower-dimensional dense sphere of unit-normed …
Study shows fast rates for inverse reinforcement learning with linear rewards.
A closed four dimensional manifold cannot possess a non-flat Ricci soliton metric with arbitrarily small -norm of the curvature. In this paper, we localize this fact in the case of shrinking Ricci solitons by proving an -regularity theorem, thus confirming a conjecture of Cheeger-Tian. As applications…
Entropic regularization is quickly emerging as a new standard in optimal transport (OT). It enables to cast the OT computation as a differentiable and unconstrained convex optimization problem, which can be efficiently solved using the Sinkhorn algorithm. However, entropy keeps the transportation plan strictly positive…
New research shows SAA can outperform SA for Wasserstein barycenters.
Learning with a primary objective, such as softmax cross entropy for classification and sequence generation, has been the norm for training deep neural networks for years. Although being a widely-adopted approach, using cross entropy as the primary objective exploits mostly the information from the ground-truth class f…
Learning in Deep Neural Networks (DNN) takes place by minimizing a non-convex high-dimensional loss function, typically by a stochastic gradient descent (SGD) strategy. The learning process is observed to be able to find good minimizers without getting stuck in local critical points, and that such minimizers are often …
Logit dynamics formula reveals self-regulation in softmax policy gradient methods.
A new method speeds up SoftMax normalization for embedding learning.
This paper tackles over-certainty in test-time adaptation models, proposing a solution to improve calibration.