This paper shows that one cannot learn the probability of rare events without imposing further structural assumptions. The event of interest is that of obtaining an outcome outside the coverage of an i.i.d. sample from a discrete distribution. The probability of this event is referred to as the "missing mass". The impo…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new method for averaging probability distributions based on optimal weak mass transport.
Novel concentration inequalities are obtained for the missing mass, i.e. the total probability mass of the outcomes not observed in the sample. We derive distribution-free deviation bounds with sublinear exponents in deviation size for missing mass and improve the results of Berend and Kontorovich (2013) and Yari Saeed…
The study optimizes distribution estimation from samples with relative entropy error, adapting to sparse distributions.
Develops a new divergence framework that combines -divergences and IPMs.
A new method calculates fractional moments using the moment-generating function.
New method for summarizing ranking distributions using consensus ranking distributions.
We are concerned with obtaining novel concentration inequalities for the missing mass, i.e. the total probability mass of the outcomes not observed in the sample. We not only derive - for the first time - distribution-free Bernstein-like deviation bounds with sublinear exponents in deviation size for missing mass, but …
CNFs learn on manifolds using PPD, improving likelihood and sample quality.
New formulae identify discrete probability laws without needing normalization constants.
This review explores entropy applications in data analysis and machine learning.
Adaptive HMC improves sampling efficiency by optimizing mass matrix.
In this paper, we are concerned with obtaining distribution-free concentration inequalities for mixture of independent Bernoulli variables that incorporate a notion of variance. Missing mass is the total probability mass associated to the outcomes that have not been seen in a given sample which is an important quantity…
A new method quantizes conditional probability measures using deep learning.
UCPO improves diversity in reinforcement learning models, maintaining high accuracy.
Classical optimal transport problem seeks a transportation map that preserves the total mass betwenn two probability distributions, requiring their mass to be the same. This may be too restrictive in certain applications such as color or shape matching, since the distributions may have arbitrary masses and/or that only…
Identifies conditions for multiple invariant probabilities in Markov kernels.
Hamiltonian Monte Carlo (HMC) is an efficient Bayesian sampling method that can make distant proposals in the parameter space by simulating a Hamiltonian dynamical system. Despite its popularity in machine learning and data science, HMC is inefficient to sample from spiky and multimodal distributions. Motivated by the …
SDE automatically recovers interpretable discrete distributions.
Estimates stationary mass and frequency from non-i.i.d. data.
This paper finds a unique partition of a sample space for estimating continuous distributions.
We propose a deep-learning approach based on generative adversarial networks (GANs) to reduce noise in weak lensing mass maps under realistic conditions. We apply image-to-image translation using conditional GANs to the mass map obtained from the first-year data of Subaru Hyper Suprime-Cam (HSC) survey. We train the co…
In this paper, we consider the problem of classification of high dimensional queries to high dimensional classes where and are discrete alphabets and the probabilistic model that relates data to the classes is known. This problem has applications …
New method calibrates classifier probabilities with guaranteed coverage.
GT estimator shows convergence for Markov samples, improving i.i.d. results.
MF-PID uses interacting samples to efficiently transport probability mass.
Paper proposes MMC to avoid high-density bias in clustering.
We revisit logistic regression and its nonlinear extensions, including multilayer feedforward neural networks, by showing that these classifiers can be viewed as converting input or higher-level features into Dempster-Shafer mass functions and aggregating them by Dempster's rule of combination. The probabilistic output…
We propose a correlated stochastic process of which the novel non-Gaussian probability mass function is constructed by exactly solving moment generating function. The calculation of cumulants and auto-correlation shows that the process is convergent and scale invariant in the large but finite number limit. We demonstra…
The paper reviews advances in estimating and understanding optimal transport maps.
A new method estimates marginal likelihood using normalizing flows.
Study shows mass distribution of random holomorphic sections follows a central limit theorem.
Estimates joint probability distribution from 1-way marginals using low-rank tensors and random projections.
In a previous analysis the problem of "zero-inflated" time data (caused by high frequency trading in the electronic order book) was handled by left-truncating the inter-arrival times. We demonstrated, using rigorous statistical methods, that the Weibull distribution describes the corresponding stochastic dynamics for a…
A framework to quantify deployment risk in ML systems, especially for rare states.
Given samples from a population of individuals belonging to different types with unknown proportions, how do we estimate the probability of discovering a new type at the -th draw? This is a classical problem in statistics, commonly referred to as the missing mass estimation problem. Recent results by Ohannes…
In this article, we classify the set of asymptotic mass-like invariants for asymptotically hyperbolic metrics. It turns out that the standard mass is just one example (but probably the most important one) among the two families of invariants we find. These invariants are attached to finite-dimensional representations o…
Positive mass theorem for asymptotically flat manifolds with non-negative distributional scalar curvature
A robust conformal method for set estimation using non-conformity scores.
By Federer and Fleming there exist at least one mass-minimizing normal current in every real-valued homology class of a Riemannian manifold. However the regularity of the mass-minimizing currents and their distributions may generally be quite complicated. In this paper we shall study how to construct nice metrics so th…
Expected centre of mass for random embeddings is constant.
Proposes a method to construct risk-neutral marginals from arbitrage-free option prices.
Study three types of uncertainty quantification for binary classification without distributional assumptions.
In this paper, we prove Lorentzian positive mass theorem for spacetimes with distributional curvature. To do so, we introduce distributional curvature and generalized Arnowitt-Deser-Misner (ADM) momentum. As an application, we discuss a junction of spacetimes.
Evidential clustering is an approach to clustering in which cluster-membership uncertainty is represented by a collection of Dempster-Shafer mass functions forming an evidential partition. In this paper, we propose to construct these mass functions by bootstrapping finite mixture models. In the first step, we compute b…
We investigate the properties of multidimensional probability distributions in the context of latent space prior distributions of implicit generative models. Our work revolves around the phenomena arising while decoding linear interpolations between two random latent vectors -- regions of latent space in close proximit…
Novel approach for estimating joint probability densities using tensor decompositions and dictionaries.
Whether you trade futures for yourself or a hedge fund, your strategy is counted. Long and short position limits make the number of unique strategies finite. Formulas of the numbers of strategies, transactions, do nothing actions are derived. A discrete distribution of actions, corresponding probability mass, cumulativ…