Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Jun 199319922001200920182026
48 results for entropy norm

Proposes a new criterion for selecting Nash equilibria considering both utility and inequality.

problem Finding a fair Nash equilibrium in group decision-making.
method Introduces entropy-norm space for geometric selection of strict Nash equilibria.
result The closest entropy-norm pair to the largest entropy-norm pair in rescaled space is the most suitable equilibrium.

Study on automorphisms of K3 and Enriques surfaces, proving entropy gaps and achirality.

problem Entropy norms and achirality of automorphisms on K3 and Enriques surfaces.
method Proves gap theorems for entropy norms and studies achirality in terms of genus-one fibrations.
result Entropy gaps and achirality results for automorphisms of K3 and Enriques surfaces.

The paper proves inequalities linking geometric norms and Thurston norms in hyperbolic 3-manifolds.

problem Inequalities linking geometric norms and Thurston norms in hyperbolic 3-manifolds.
method Analyzes geometric L2L^2-norms, Thurston norms, and Lipschitz maps to prove inequalities.
result Proves an inequality between geometric L2L^2-norm and Thurston norm, qualitatively sharp.

This paper studies neural networks with bounded norms to avoid the curse of dimensionality.

problem The curse of dimensionality in approximating functions by neural networks.
method Investigates over-parameterized two-layer neural networks with norm constraints in RKHS.
result Improved sample complexity and generalization bounds for neural networks with bounded norms.

Characterizes sample complexity for outcome indistinguishability in machine learning.

problem Outcome indistinguishability in machine learning, focusing on distinguishers and predictors.
method Sample complexity characterized by metric entropy of predictor and distinguisher classes, using dual Minkowski norms.
result Equivalence and tightness of sample complexity characterizations in distribution-specific and distribution-free settings.

Simple regional perturbations maintain model transferability while reducing adversarial example distortion.

problem Comparing efficacy of regional adversarial attacks without complex methods.
method Developed a simple regional adversarial perturbation attack using cross-entropy sign.
result Localized adversarial examples require significantly less LpL_p norm distortion compared to non-local counterparts.

We consider a hyperbolic surface bundle over the circle with the smallest known volume among hyperbolic manifolds having 3 cusps, so called "the magic manifold". We compute the entropy function on the fiber face of the unit ball with respect to the Thurston norm, determine homology classes whose representatives are gen…

2008-12-25abs ↗pdf ↗

Sharp lower bounds on shallow neural networks' approximation rates are derived.

problem The efficiency of shallow neural networks in approximating functions.
method Lower bounding the L2L^2-metric entropy and Kolmogorov nn-widths of the convex hull of neural network basis functions.
result Sharp lower bounds on the approximation rates for shallow neural networks are provided.

This paper provides a general result on controlling local Rademacher complexities, which captures in an elegant form to relate the complexities with constraint on the expected norm to the corresponding ones with constraint on the empirical norm. This result is convenient to apply in real applications and could yield re…

2015-10-06abs ↗pdf ↗

We consider the minimum error entropy (MEE) criterion and an empirical risk minimization learning algorithm in a regression setting. A learning theory approach is presented for this MEE algorithm and explicit error bounds are provided in terms of the approximation ability and capacity of the involved hypothesis space w…

2012-08-03abs ↗pdf ↗

The relaxed maximum entropy problem is concerned with finding a probability distribution on a finite set that minimizes the relative entropy to a given prior distribution, while satisfying relaxed max-norm constraints with respect to a third observed multinomial distribution. We study the entire relaxation path for thi…

2013-11-07abs ↗pdf ↗

New microcanonical models approximate non-Gaussian processes with long-range correlations.

problem Approximating non-Gaussian stationary processes with long-range correlations.
method Microcanonical models conditioned by energy vector, gradient descent, multiscale energy vectors.
result Microcanonical gradient descent processes converge and capture sparsity.

Paper explains neural collapse in neural networks using a new model.

problem Understanding neural collapse in neural networks during training.
method Introducing the unconstrained layer-peeled model (ULPM) to prove gradient flow convergence to critical points of a minimum-norm separation problem.
result Proves that all critical points are strict saddle points except the global minimizers exhibiting neural collapse.

New method prevents entropy collapse in Transformer training, leading to more stable and robust models.

problem Training instability in Transformers, especially in attention layers.
method Spectral normalization with a learned scalar to prevent entropy collapse.
result Prevents entropy collapse, leading to more stable training.

In this paper, we consider low rank matrix estimation using either matrix-version Dantzig Selector A^λd\hat{A}_λ^d or matrix-version LASSO estimator A^λL\hat{A}_λ^L. We consider sub-Gaussian measurements, i.e.i.e., the measurements X1,,XnRm×mX_1,\ldots,X_n\in\mathbb{R}^{m\times m} have i.i.d.i.i.d. sub-Gaussian entries. Suppose $\textrm…

2014-03-25abs ↗pdf ↗

Lower bound shows no acceleration for specific convex optimization class.

problem Proving lower bounds for convergence rates of convex optimization methods.
method Proving Ω(L/T)Ω(L/T) lower bound for minimization of convex and LL-smooth functions relative to negative entropy.
result Mirror descent is optimal up to a logarithmic factor in the class of functions considered.

We consider the problem of online nonparametric regression with arbitrary deterministic sequences. Using ideas from the chaining technique, we design an algorithm that achieves a Dudley-type regret bound similar to the one obtained in a non-constructive fashion by Rakhlin and Sridharan (2014). Our regret bound is expre…

2015-02-26abs ↗pdf ↗

New scalable algorithm for non-negative linear regression with entropy-regularized OT loss.

problem Generalizing task-specific linear models to broader applications.
method Sinkhorn-like scaling iterations for convex penalty and datafit terms.
result Simple multiplicative updates for various penalty and datafit terms.

This paper proves a generalization bound for complex-valued neural networks scaling with spectral complexity.

problem Ensuring the performance of complex-valued neural networks on unseen data.
method Theoretical derivation using Maurey Sparsification Lemma and Dudley Entropy Integral, empirical validation on various datasets.
result The spectral complexity of weight matrices is a significant factor in the generalization ability of complex-valued neural networks.

The density matrices are positively semi-definite Hermitian matrices of unit trace that describe the state of a quantum system. The goal of the paper is to develop minimax lower bounds on error rates of estimation of low rank density matrices in trace regression models used in quantum state tomography (in particular, i…

2015-07-17abs ↗pdf ↗

Gradient descent biases linear models in next-token prediction towards data entropy.

problem Optimization bias in next-token prediction models.
method Analysis of gradient descent on linear models with sparse conditional distributions.
result Gradient descent selects parameters that equate token logits differences to log-odds in the data subspace.

Recent contributions have framed linear system identification as a nonparametric regularized inverse problem. Relying on 2\ell_2-type regularization which accounts for the stability and smoothness of the impulse response to be estimated, these approaches have been shown to be competitive w.r.t classical parametric met…

2015-08-12abs ↗pdf ↗

Study bounds on kernel function entropy for finite measures.

problem Investigate bounds on the ε-entropy of kernel classes.
method Sharp upper and lower bounds for p in [1, +∞] derived from eigenvalue behavior and Mercer series convergence.
result Proves tighter bounds for general kernels compared to previous work.

A method for learning from unlabeled time-series data using temporal smoothing and entropy maximization.

problem Learning from unlabeled time-series data efficiently and accurately.
method Training a feedforward neural network with two objectives: temporal smoothing and entropy maximization.
result The method extracts slowly evolving information from time-series data, filtering out noise.

We study the problem of reconstructing an unknown matrix M of rank r and dimension d using O(rd poly log d) Pauli measurements. This has applications in quantum state tomography, and is a non-commutative analogue of a well-known problem in compressed sensing: recovering a sparse vector from a few of its Fourier coeffic…

2011-03-14abs ↗pdf ↗

Study on convergence of exponential probability measures with applications to maximum entropy models and SGLD.

problem Characterizing the limit of probability measures with exponential densities as temperature approaches zero.
method Quantitative bounds on Wasserstein distance using geometric measure theory tools.
result Established quantitative convergence results for norm-like potentials under invertibility conditions.

The paper analyzes how well classes are separated in neural network feature space.

problem Understanding class separability in neural network feature space.
method Theoretical analysis of intra-class and inter-class distances in feature space.
result A lower bound for the probability of inter-class distance being greater than intra-class distance as a function of loss value.

Study shows fast rates for inverse reinforcement learning with linear rewards.

problem Entropy-regularized min-max inverse reinforcement learning in finite-horizon MDPs.
method Structural and statistical analysis of Min-Max-IRL with pseudo-self-concordance.
result Both trajectory-level KL divergence and parameter error decay at O(n1)\mathcal{O}(n^{-1}).

The paper explores how neural networks avoid overfitting by learning from flat minima.

problem How to prevent overfitting in neural networks by learning from flat minima.
method Study of one- and two-layer neural network models, derivation of algorithms focusing on wide flat regions.
result Wide flat minima coexist with narrower minima and critical points, and are associated with good minimizers.

Cognitive abstraction metrics predict deep neural networks' generalization ability.

problem Understanding deep neural networks' generalization ability.
method Defined Cognitive Neural Activation (CNA) metric based on information complexity and activation patterns.
result CNA is highly predictive of generalization ability, outperforming other metrics.

New research shows SAA can outperform SA for Wasserstein barycenters.

problem Optimizing Wasserstein barycenters with entropy regularization.
method Comparison of Stochastic Approximation (SA) and Sample Average Approximation (SAA) for large-scale problems.
result SAA can be more efficient than SA for Wasserstein barycenters, especially in large-scale settings.

A new method speeds up SoftMax normalization for embedding learning.

problem Efficiently learning distributed representations with SoftMax normalization.
method Proposes a linear-time heuristic approximation for mSoftMax(XYT){ m SoftMax}(XY^T), optimizing cross entropy.
result Achieves higher or comparable accuracy to existing methods with lower computational time.