New statistics are introduced that maintain the Fisher metric structure closely, akin to sufficient statistics.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper introduces data-dependent SSP for private linear and logistic regression.
New method uses sufficient statistics to infer causal relationships from observational data.
New statistical theory explains contrastive learning effectiveness.
Information geometry provides a geometric approach to families of statistical models. The key geometric structures are the Fisher quadratic form and the Amari-Chentsov tensor. In statistics, the notion of sufficient statistic expresses the criterion for passing from one model to another without loss of information. Thi…
One of the most fundamental questions one can ask about a pair of random variables X and Y is the value of their mutual information. Unfortunately, this task is often stymied by the extremely large dimension of the variables. We might hope to replace each variable by a lower-dimensional representation that preserves th…
Boosting improves data fitting while maintaining fairness guarantees.
Reduces IB problem to a simpler, lower-dimensional problem.
We introduce Minimal Achievable Sufficient Statistic (MASS) Learning, a training method for machine learning models that attempts to produce minimal sufficient statistics with respect to a class of functions (e.g. deep networks) being optimized over. In deriving MASS Learning, we also introduce Conserved Differential I…
We uncover a fairly general principle in online learning: If regret can be (approximately) expressed as a function of certain "sufficient statistics" for the data sequence, then there exists a special Burkholder function that 1) can be used algorithmically to achieve the regret bound and 2) only depends on these suffic…
We show that, for an affine submersion with horizontal distribution, is a statistical manifold with the metric and connection induced from the statistical manifold . The concept of conformal submersion with horizontal distribution is introduced, which i…
This paper introduces SS-MAMP to address convergence issues in AMP algorithms.
The principal support vector machines method (Li et al., 2011) is a powerful tool for sufficient dimension reduction that replaces original predictors with their low-dimensional linear combinations without loss of information. However, the computational burden of the principal support vector machines method constrains …
Efficiently learns Ising model parameters with limited statistics.
This paper provides a method for noise-calibrated inference from DP synthetic data.
We investigate the problem of learning discrete, undirected graphical models in a differentially private way. We show that the approach of releasing noisy sufficient statistics using the Laplace mechanism achieves a good trade-off between privacy, utility, and practicality. A naive learning algorithm that uses the nois…
Neural networks help create summary statistics for complex models.
Paper develops a theory explaining contrastive pre-training for multimodal AI.
This paper is about metric data structures in high-dimensional or non-Euclidean space that permit cached sufficient statistics accelerations of learning algorithms. It has recently been shown that for less than about 10 dimensions, decorating kd-trees with additional "cached sufficient statistics" such as first and sec…
Chentsov's theorem characterizes the Fisher information metric on statistical models as essentially the only Riemannian metric that is invariant under sufficient statistics. This implies that each statistical model is naturally equipped with a geometry, so Chentsov's theorem explains why many statistical properties can…
In this note we prove certain necessary and sufficient conditions for the existence of an embedding of statistical manifolds. In particular, we prove that any compact smooth ( resp.) statistical manifold can be embedded into the space of probability measures on a finite set. As a result, we get an answer to the La…
This paper studies the geometry of immersions into statistical manifolds. A necessary and sufficient condition is obtained for statistical manifold structures to be dual to each other for a non-degenerate equiaffine immersion. Then we obtain conditions for realizing an n-dimensional statistical manifold in an (n+1)-dim…
The minimum message length principle is an information theoretic criterion that links data compression with statistical inference. This paper studies the strict minimum message length (SMML) estimator for -dimensional exponential families with continuous sufficient statistics, for all . The partition of an …
The φ-sectional curvature of statistical structures on almost contact metric manifolds is always non-positive.
We develop amortized population Gibbs (APG) samplers, a class of scalable methods that frames structured variational inference as adaptive importance sampling. APG samplers construct high-dimensional proposals by iterating over updates to lower-dimensional blocks of variables. We train each conditional proposal by mini…
Transformers encode latent distributions in text, improving performance in out-of-distribution cases.
A new principle and method improve out-of-distribution detection in generative models.
We investigate a generic problem of learning pairwise exponential family graphical models with pairwise sufficient statistics defined by a global mapping function, e.g., Mercer kernels. This subclass of pairwise graphical models allow us to flexibly capture complex interactions among variables beyond pairwise product. …
A condition for a statistical manifold to have an equiaffine structure is studied. The facts that dual flatness and conjugate symmetry of a statistical manifold are sufficient conditions for a statistical manifold to have an equiaffine structure were obtained in [2] and [3]. In this paper, a fact that a statistical man…
Develops algorithms to balance personalization and statistical power in mobile health studies.
We propose a novel approach for density estimation with exponential families for the case when the true density may not fall within the chosen family. Our approach augments the sufficient statistics with features designed to accumulate probability mass in the neighborhood of the observed points, resulting in a non-para…
Neural networks simplify SDR in regression tasks.
This work uses statistical bootstrapping to provide accurate confidence intervals for policy value in reinforcement learning.
Conditions for statistical structures on manifolds derived from solitons.
New method for separating mixed signals with nonlinear functions.
FSRL balances fairness and sufficiency in learning representations.
Paper compares hard and soft EM for BN learning from incomplete data.
We define and study the statistical models in exponential family form whose sufficient statistics are the degree distributions and the bi-degree distributions of undirected labelled simple graphs. Graphs that are constrained by the joint degree distributions are called -graphs in the computer science literature and…
We analyse the learning performance of Distributed Gradient Descent in the context of multi-agent decentralised non-parametric regression with the square loss function when i.i.d. samples are assigned to agents. We show that if agents hold sufficiently many samples with respect to the network size, then Distributed Gra…
New method needed for class prior estimation when covariates are reduced.
Geometric regularisation improves statistical models by avoiding degeneracy loci.
This paper develops embeddings that preserve likelihood-based statistical inference.
The claim arrival process to an insurance company is modeled by a compound Poisson process whose intensity and/or jump size distribution changes at an unobservable time with a known distribution. It is in the insurance company's interest to detect the change time as soon as possible in order to re-evaluate a new fair v…
Generalizes data thinning for various distributions.
Study on forecasting methods and their causal implications.
Paper introduces a new gradient statistic to improve deep learning convergence.
FF algorithm uses goodness as a likelihood-ratio test for scalar normalization.
Information geometry offers new tools for statistical analysis.