This paper shows that one cannot learn the probability of rare events without imposing further structural assumptions. The event of interest is that of obtaining an outcome outside the coverage of an i.i.d. sample from a discrete distribution. The probability of this event is referred to as the "missing mass". The impo…
A new method for averaging probability distributions based on optimal weak mass transport.
problem Averaging probability distributions in a geometric way.
method Weak barycenters based on optimal weak mass transport.
result Extracts common geometric information shared by all input distributions.
Novel concentration inequalities are obtained for the missing mass, i.e. the total probability mass of the outcomes not observed in the sample. We derive distribution-free deviation bounds with sublinear exponents in deviation size for missing mass and improve the results of Berend and Kontorovich (2013) and Yari Saeed…
The study optimizes distribution estimation from samples with relative entropy error, adapting to sparse distributions.
problem Estimating discrete distributions with high-probability accuracy in relative entropy.
method Analysis of Laplace estimator and confidence-dependent smoothing techniques, including data-dependent smoothing.
result Optimal high-probability risk bounds for various estimators, including a new data-dependent smoothing method.
Develops a new divergence framework that combines f-divergences and IPMs.
problem Comparing distributions that are not absolutely continuous.
method Introduces (f,Γ)-divergences as a two-stage mass-redistribution/mass-transport process. result Improves estimation, learning, and uncertainty quantification in GANs for heavy-tailed distributions.
A new method calculates fractional moments using the moment-generating function.
problem Computing fractional moments from probability densities.
method Integral framework based on moment-generating function.
result Exact integral expressions for various types of moments.
New method for summarizing ranking distributions using consensus ranking distributions.
problem Summarizing ranking distributions efficiently and accurately.
method Introducing consensus ranking distributions and a top-down tree-structured statistical algorithm.
result Optimal distortion can be expressed as a function of pairwise probabilities, enabling efficient learning methods.
We are concerned with obtaining novel concentration inequalities for the missing mass, i.e. the total probability mass of the outcomes not observed in the sample. We not only derive - for the first time - distribution-free Bernstein-like deviation bounds with sublinear exponents in deviation size for missing mass, but …
CNFs learn on manifolds using PPD, improving likelihood and sample quality.
problem Training CNFs on manifolds efficiently and accurately.
method Minimizing PPD, a novel divergence, to train CNFs on manifolds.
result CNFs trained with PPD achieve state-of-the-art results on manifold benchmarks.
New formulae identify discrete probability laws without needing normalization constants.
problem Characterizing non-normalized discrete probability distributions.
method Derive explicit formulae for mass functions using Stein's method.
result Developed tools for solving statistical problems without normalization constants.
This review explores entropy applications in data analysis and machine learning.
problem Characterizing probability mass distributions in data analysis and machine learning.
method Review of various entropy types and their applications.
result Entropy's versatility in data analysis and machine learning.
Adaptive HMC improves sampling efficiency by optimizing mass matrix.
problem Inefficient HMC performance due to mass matrix choice.
method Gradient-based adaptation of mass matrix to maximize proposal entropy.
result Adaptation method outperforms HMC variants by optimizing mass matrix.
In this paper, we are concerned with obtaining distribution-free concentration inequalities for mixture of independent Bernoulli variables that incorporate a notion of variance. Missing mass is the total probability mass associated to the outcomes that have not been seen in a given sample which is an important quantity…
A new method quantizes conditional probability measures using deep learning.
problem Quantizing conditional probability measures efficiently.
method DCMQ method using Huber-energy kernel and deep neural network.
result Promising results on various examples.
UCPO improves diversity in reinforcement learning models, maintaining high accuracy.
problem RLVR objectives often lead to diversity collapse, reducing coverage of correct solutions.
method UCPO adds a conditional uniformity penalty to GRPO, redistributing probability mass.
result UCPO improves Pass@K and diversity while maintaining competitive Pass@1 accuracy.
Identifies conditions for multiple invariant probabilities in Markov kernels.
problem Global irreducibility and recurrence do not guarantee uniqueness of invariant probabilities.
method Uses Jordan decomposition of the difference of two invariant probabilities.
result A Markov kernel has more than one invariant probability if and only if it admits a visible absorbing decomposition.
Hamiltonian Monte Carlo (HMC) is an efficient Bayesian sampling method that can make distant proposals in the parameter space by simulating a Hamiltonian dynamical system. Despite its popularity in machine learning and data science, HMC is inefficient to sample from spiky and multimodal distributions. Motivated by the …
SDE automatically recovers interpretable discrete distributions.
problem Limited interpretable discrete probability laws.
method Unsupervised framework using symbolic density estimation.
result Accurately recovers interpretable discrete distributions.
Estimates stationary mass and frequency from non-i.i.d. data.
problem Estimating stationary mass and frequency from non-i.i.d. data.
method Combines plug-in estimator with WingIt modification for exponentially α-mixing processes. result Universal consistency in n for total variation distance estimation. This paper finds a unique partition of a sample space for estimating continuous distributions.
problem Estimating continuous probability distributions from finite samples.
method Equal-probability partition of the sample space using order statistics.
result The partition yields an entropy of log2(N+1) bits, providing a discrete entropy estimate.
We propose a deep-learning approach based on generative adversarial networks (GANs) to reduce noise in weak lensing mass maps under realistic conditions. We apply image-to-image translation using conditional GANs to the mass map obtained from the first-year data of Subaru Hyper Suprime-Cam (HSC) survey. We train the co…
In this paper, we consider the problem of classification of M high dimensional queries y1,⋯,yM∈BS to N high dimensional classes x1,⋯,xN∈AS where A and B are discrete alphabets and the probabilistic model that relates data to the classes P(x,y) is known. This problem has applications …
New method calibrates classifier probabilities with guaranteed coverage.
problem Inaccurate probability estimates by classifiers in high-risk applications.
method Adaptive temperature scaling algorithm for conformal prediction.
result Improves calibration error measures and standard metrics across various tasks.
GT estimator shows convergence for Markov samples, improving i.i.d. results.
problem Estimating missing mass in Markov samples.
method Analyzed convergence of Good-Turing estimator for Markov samples, considering spectral properties of transition matrices.
result The convergence of the GT estimator for Markov samples depends on the spectral properties of the transition matrices, leading to a new minimax rate of 1/(nβ5) for rank-2 Markov chains. New algorithms solve partial optimal transport problems for applications like PU learning.
problem Optimal transport constraints on equal mass distributions limit applicability.
method Developed exact algorithms for partial Wasserstein and Gromov-Wasserstein problems.
result Partial Wasserstein metrics show effectiveness in positive-unlabeled learning.
MF-PID uses interacting samples to efficiently transport probability mass.
problem Efficiently transporting probability mass in generative models.
method Introducing Mean-Field Path-Integral Diffusion (MF-PID) where samples become interacting agents.
result MF-PID achieves 19-24% reductions in control energy for demand-response control of energy systems.
Paper proposes MMC to avoid high-density bias in clustering.
problem High-density bias in density-based clustering.
method Introduces mass distribution as a better foundation for clustering, proposing mass-maximization clustering (MMC).
result MMC avoids high-density bias and discovers clusters of arbitrary shapes, sizes, and densities.
We revisit logistic regression and its nonlinear extensions, including multilayer feedforward neural networks, by showing that these classifiers can be viewed as converting input or higher-level features into Dempster-Shafer mass functions and aggregating them by Dempster's rule of combination. The probabilistic output…
We propose a correlated stochastic process of which the novel non-Gaussian probability mass function is constructed by exactly solving moment generating function. The calculation of cumulants and auto-correlation shows that the process is convergent and scale invariant in the large but finite number limit. We demonstra…
The paper reviews advances in estimating and understanding optimal transport maps.
problem Estimating and understanding optimal transport maps from samples.
method Recent advances in statistical inference for optimal transport maps.
result Developed limit theorems for the optimal transport map using samples.
A new method estimates marginal likelihood using normalizing flows.
problem Estimating marginal likelihood in Bayesian model selection.
method Learned harmonic mean estimator using normalizing flows.
result Normalizing flows avoid the exploding variance problem.
Study shows mass distribution of random holomorphic sections follows a central limit theorem.
problem Understanding mass distribution of random holomorphic sections.
method Proved a central limit theorem for mass distribution of random holomorphic sections associated with positive line bundles.
result Almost every sequence of random holomorphic sections exhibits quantum ergodicity.
Estimates joint probability distribution from 1-way marginals using low-rank tensors and random projections.
problem Nonparametric estimation of joint probability mass function (PMF) from limited data.
method Low-rank tensor decomposition and random projections to link data to PMF estimation.
result Estimates joint density from 1-way marginals using transformed space and novel algorithm.
In a previous analysis the problem of "zero-inflated" time data (caused by high frequency trading in the electronic order book) was handled by left-truncating the inter-arrival times. We demonstrated, using rigorous statistical methods, that the Weibull distribution describes the corresponding stochastic dynamics for a…
A framework to quantify deployment risk in ML systems, especially for rare states.
problem Under-supported rare states in ML models lead to unreliable performance in unseen data.
method Blind-Spot Mass (B_n(tau)) using Good-Turing unseen-species estimation.
result Identifies and quantifies the risk of under-supported states in ML models.
Given n samples from a population of individuals belonging to different types with unknown proportions, how do we estimate the probability of discovering a new type at the (n+1)-th draw? This is a classical problem in statistics, commonly referred to as the missing mass estimation problem. Recent results by Ohannes…
In this article, we classify the set of asymptotic mass-like invariants for asymptotically hyperbolic metrics. It turns out that the standard mass is just one example (but probably the most important one) among the two families of invariants we find. These invariants are attached to finite-dimensional representations o…
Positive mass theorem for asymptotically flat manifolds with non-negative distributional scalar curvature
problem Positive mass theorem
method Ricci flow smoothing
result Asymptotically flat manifolds with non-negative ADM mass
A robust conformal method for set estimation using non-conformity scores.
problem Lack of robustness in standard conformal prediction methods for outliers or heavy tails.
method Robust conformal method based on non-conformity score defined as half-mass radius.
result Empirical conformal regions converge to robust population central set.
By Federer and Fleming there exist at least one mass-minimizing normal current in every real-valued homology class of a Riemannian manifold. However the regularity of the mass-minimizing currents and their distributions may generally be quite complicated. In this paper we shall study how to construct nice metrics so th…
Expected centre of mass for random embeddings is constant.
problem Understanding the expected centre of mass for random embeddings.
method Analyzing the Haar measure and Gaussian unitary ensemble on SL(N, C).
result The expectation of the centre of mass is a constant multiple of the identity matrix.
Proposes a method to construct risk-neutral marginals from arbitrage-free option prices.
problem Lack of risk-neutral marginals that are free of arbitrage and easy to use.
method Explicit construction of risk-neutral marginals from discrete arbitrage-free option prices.
result Explicit construction guarantees risk-neutral marginals free of butterfly and calendar arbitrage.
Study three types of uncertainty quantification for binary classification without distributional assumptions.
problem Uncertainty quantification for binary classification in a distribution-free setting.
method Established theorems connecting calibration, confidence intervals, and prediction sets for score-based classifiers.
result Distribution-free calibration is only possible using scoring functions that partition feature space into countably many sets.
In this paper, we prove Lorentzian positive mass theorem for spacetimes with distributional curvature. To do so, we introduce distributional curvature and generalized Arnowitt-Deser-Misner (ADM) momentum. As an application, we discuss a junction of spacetimes.
Evidential clustering is an approach to clustering in which cluster-membership uncertainty is represented by a collection of Dempster-Shafer mass functions forming an evidential partition. In this paper, we propose to construct these mass functions by bootstrapping finite mixture models. In the first step, we compute b…
We investigate the properties of multidimensional probability distributions in the context of latent space prior distributions of implicit generative models. Our work revolves around the phenomena arising while decoding linear interpolations between two random latent vectors -- regions of latent space in close proximit…
Novel approach for estimating joint probability densities using tensor decompositions and dictionaries.
problem Estimating joint probability densities of mixed discrete and continuous variables.
method Low-rank tensor decomposition combined with dictionary learning.
result Better classification and lower error rates compared to existing methods.
Method generates joint posterior samples of source and foreground mass distributions for gravitational lensing.
problem Challenging inference problem for high-resolution, high signal-to-noise ratio gravitational lensing.
method Combines diffusion-based generative modeling and recurrent inference machines.
result Can model realistic gravitational lensing simulations down to the noise level.