Paper introduces kernel deformed exponential families for sparse continuous attention.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Efficient method for learning continuous exponential families beyond Gaussian.
New bounds for score matching in polynomial exponential families.
Paper extends sparse alternatives to softmax for continuous domains, enabling efficient attention mechanisms.
This paper develops sparse alternatives to continuous distributions, including new types of Gaussians and attention mechanisms.
This paper analyzes VAE approximation errors in conditional exponential families.
Chentsov's theorem characterizes the Fisher information metric on statistical models as essentially the only Riemannian metric that is invariant under sufficient statistics. This implies that each statistical model is naturally equipped with a geometry, so Chentsov's theorem explains why many statistical properties can…
Paper interprets DNNs using RG for exponential family data.
BART is extended to handle various response variables.
Conjugate pairs of distributions over infinite dimensional spaces are prominent in statistical learning theory, particularly due to the widespread adoption of Bayesian nonparametric methodologies for a host of models and applications. Much of the existing literature in the learning community focuses on processes posses…
The minimum message length principle is an information theoretic criterion that links data compression with statistical inference. This paper studies the strict minimum message length (SMML) estimator for -dimensional exponential families with continuous sufficient statistics, for all . The partition of an …
Simplex-valued data appear throughout statistics and machine learning, for example in the context of transfer learning and compression of deep networks. Existing models for this class of data rely on the Dirichlet distribution or other related loss functions; here we show these standard choices suffer systematically fr…
Exponential families and mixture families are parametric probability models that can be geometrically studied as smooth statistical manifolds with respect to any statistical divergence like the Kullback-Leibler (KL) divergence or the Hellinger divergence. When equipping a statistical manifold with the KL divergence, th…
Boosting improves data fitting while maintaining fairness guarantees.
Proves convergence of mean curvature flow on cylinders with unique continuation.
Exponential family distributions are highly useful in machine learning since their calculation can be performed efficiently through natural parameters. The exponential family has recently been extended to the t-exponential family, which contains Student-t distributions as family members and thus allows us to handle noi…
The study explores generalized divergences and exponential families with a focus on sufficient conditions and laws of large numbers.
Correspondence found between exponential families and affine Grassmannians.
We analyze a plug-in estimator for a large class of integral functionals of one or more continuous probability densities. This class includes important families of entropy, divergence, mutual information, and their conditional versions. For densities on the -dimensional unit cube that lie in a -Hölder s…
A method is given for calculating the strict minimum message length (SMML) estimator for 1-dimensional exponential families with continuous sufficient statistics. A set of equations are found that the cut-points of the SMML estimator must satisfy. These equations can be solved using Newton's method and this app…
We provide a classification of graphical models according to their representation as subfamilies of exponential families. Undirected graphical models with no hidden variables are linear exponential families (LEFs), directed acyclic graphical models and chain graphs with no hidden variables, including Bayesian networks …
Constructing exponential families from statistical manifolds.
We consider the matrix completion problem of recovering a structured matrix from noisy and partial measurements. Recent works have proposed tractable estimators with strong statistical guarantees for the case where the underlying matrix is low--rank, and the measurements consist of a subset, either of the exact individ…
Score matching offers efficient estimation for certain distributions.
EFA extends self-attention to handle mixed data types and dynamic relevance.
Thompson Sampling has been demonstrated in many complex bandit models, however the theoretical guarantees available for the parametric multi-armed bandit are still limited to the Bernoulli case. Here we extend them by proving asymptotic optimality of the algorithm using the Jeffreys prior for 1-dimensional exponential …
Moment polytope of toric exponential families is a projection of a simplex.
New Thompson sampling algorithm reduces regret for exponential family bandits.
We study online learning under logarithmic loss with regular parametric models. Hedayati and Bartlett (2012b) showed that a Bayesian prediction strategy with Jeffreys prior and sequential normalized maximum likelihood (SNML) coincide and are optimal if and only if the latter is exchangeable, and if and only if the opti…
A tractable pseudo-metric for non-parametric distributions via SPD geometry.
We propose a novel approach for density estimation with exponential families for the case when the true density may not fall within the chosen family. Our approach augments the sufficient statistics with features designed to accumulate probability mass in the neighborhood of the observed points, resulting in a non-para…
The versatility of exponential families, along with their attendant convexity properties, make them a popular and effective statistical model. A central issue is learning these models in high-dimensions, such as when there is some sparsity pattern of the optimal parameter. This work characterizes a certain strong conve…
New insights into natural exponential families improve regret bounds for bandit problems.
Extends likelihood ratio exponential families to analyze various optimization methods.
Exact discrete mechanics for nonholonomic systems defined.
Proves properties of sub-Riemannian exponential map, showing it's not injective.
Maximum likelihood learning with exponential families leads to moment-matching of the sufficient statistics, a classic result. This can be generalized to conditional exponential families and/or when there are hidden data. This document gives a first-principles explanation of these generalized moment-matching conditions…
EFDA extends LDA to non-Gaussian models using exponential families.
We study optimal solutions to an abstract optimization problem for measures, which is a generalization of classical variational problems in information theory and statistical physics. In the classical problems, information and relative entropy are defined using the Kullback-Leibler divergence, and for this reason optim…
Exponential family extensions of principal component analysis (EPCA) have received a considerable amount of attention in recent years, demonstrating the growing need for basic modeling tools that do not assume the squared loss or Gaussian distribution. We extend the EPCA model toolbox by presenting the first exponentia…
We consider a family of learning strategies for online optimization problems that evolve in continuous time and we show that they lead to no regret. From a more traditional, discrete-time viewpoint, this continuous-time approach allows us to derive the no-regret properties of a large class of discrete-time algorithms i…
New hyperbolic manifolds show exponential homology torsion growth.
Improves variational inference for sparse models using mixtures of exponential families.
SMRL uses score matching for efficient RL with exponential family models.
We prove the existence of an abundance of new Einstein metrics on odd dimensional spheres including exotic spheres, many of them depending on continuous parameters. The number of families as well as the number of parameter grows double exponentially with the dimension. Our method of proof uses Brieskorn-Pham singularit…
AdjointDEIS simplifies diffusion model optimization.
CDEFs reduce model complexity and uncover time correlations.
Deep equilibrium models estimate latent variables from data.