The study optimizes bounds for comparing training and population loss.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We are interested in comparing probability distributions defined on Riemannian manifold. The traditional approach to study a distribution relies on locating its mean point and finding the dispersion about that point. On a general manifold however, even if two distributions are sufficiently concentrated and have unique …
Paper addresses fault-tolerance in distributed machine learning with stochastic gradient descent.
The Poisson distribution has been widely studied and used for modeling univariate count-valued data. Multivariate generalizations of the Poisson distribution that permit dependencies, however, have been far less popular. Yet, real-world high-dimensional count-valued data found in word counts, genomics, and crime statis…
Shortage of labeled data has been holding the surge of deep learning in healthcare back, as sample sizes are often small, patient information cannot be shared openly, and multi-center collaborative studies are a burden to set up. Distributed machine learning methods promise to mitigate these problems. We argue for a sp…
There is emerging interest in performing regression between distributions. In contrast to prediction on single instances, these machine learning methods can be useful for population-based studies or on problems that are inherently statistical in nature. The recently proposed distribution regression network (DRN) has sh…
Study compares multivariate scoring rules for distribution forecasts.
Omega ratio, defined as the probability-weighted ratio of gains over losses at a given level of expected return, has been advocated as a better performance indicator compared to Sharpe and Sortino ratio as it depends on the full return distribution and hence encapsulates all information about risk and return. We comput…
Study compares synthetic and distributional Ricci curvature bounds.
We propose a new Integral Probability Metric (IPM) between distributions: the Sobolev IPM. The Sobolev IPM compares the mean discrepancy of two distributions for functions (critic) restricted to a Sobolev ball defined with respect to a dominant measure . We show that the Sobolev IPM compares two distributions in hig…
Paper establishes sufficient condition for comparing linear combinations of infinite-mean risks.
This work bridges outlier and drift detection by comparing inputs to a part of the reference distribution.
Since their introduction a year ago, distributional approaches to reinforcement learning (distributional RL) have produced strong results relative to the standard approach which models expected values (expected RL). However, aside from convergence guarantees, there have been few theoretical results investigating the re…
We identify a fundamental problem in policy gradient-based methods in continuous control. As policy gradient methods require the agent's underlying probability distribution, they limit policy representation to parametric distribution classes. We show that optimizing over such sets results in local movement in the actio…
UAPCA projects uncertain data to low dimensions using GMMs.
We discuss some basic concepts of semi-Riemannian geometry in low-regularity situations. In particular, we compare the settings of (linear) distributional geometry in the sense of L. Schwartz and nonlinear distributional geometry in the sense of J.F. Colombeau.
We undertake a systematic comparison between implied volatility, as represented by VIX (new methodology) and VXO (old methodology), and realized volatility. We compare visually and statistically distributions of realized and implied variance (volatility squared) and study the distribution of their ratio. We find that t…
Distributed Lion optimizes large model training by reducing communication costs.
Paper compares generative and discriminative models in uncertainty quantification.
We introduce blockchains and distributed ledgers and describe their potential applications to money and banking. The analysis compares public and private ledgers and outlines the suitability of various types of ledgers for different purposes. Furthermore, a few historical prototypes of blockchains and distributed ledge…
Robust forecast framework reduces distribution error by 63%.
Study compares Bitcoin and Ethereum tail behavior using Q-Q plots.
This note presents an operational measure of fat-tailedness for univariate probability distributions, in where 0 is maximally thin-tailed (Gaussian) and 1 is maximally fat-tailed. Among others,1) it helps assess the sample size needed to establish a comparative needed for statistical significance, 2) allows…
The ability to represent and compare machine learning models is crucial in order to quantify subtle model changes, evaluate generative models, and gather insights on neural network architectures. Existing techniques for comparing data distributions focus on global data properties such as mean and covariance; in that se…
This article reviews and compares various methods for estimating conditional distributions.
Multi-instance data, in which each object (bag) contains a collection of instances, are widespread in machine learning, computer vision, bioinformatics, signal processing, and social sciences. We present a maximum entropy (ME) framework for learning from multi-instance data. In this approach each bag is represented as …
Paper introduces S3W distance for spherical probability distributions.
AF improves sampling from high-dimensional, multi-modal distributions.
The study applies spatial density models to mobile node movements using Möbius distributions.
New distributed clustering algorithms show resilience to initialization issues.
Study compares data-driven vs model-based MRS quantification strategies, focusing on resilience to out-of-distribution effects.
New framework improves learning across multiple distributions.
This paper applies the Extreme-Value (EV) Generalised Pareto distribution to the extreme tails of the return distributions for the S&P500, FT100, DAX, Hang Seng, and Nikkei225 futures contracts. It then uses tail estimators from these contracts to estimate spectral risk measures, which are coherent risk measures that r…
A new Wasserstein distance method for comparing incomparable distributions.
This paper introduces a new method to compare collections of distributions on manifolds and graphs.
Proposes diffusion models using mixed Gaussian priors for better data representation.
GNNS uses graph neural networks to efficiently estimate subgraph frequency distributions.
In this paper, we focus on weakly supervised learning with noisy training data for both classification and regression problems.We assume that the training outputs are collected from a mixture of a target and correlated noise distributions.Our proposed method simultaneously estimates the target distribution and the qual…
Study compares and accelerates deep learning for extreme events modeling.
Locus scores predictions for risk, reducing large-loss events.
This paper applies an AR(1)-GARCH (1, 1) process to detail the conditional distributions of the return distributions for the S&P500, FT100, DAX, Hang Seng, and Nikkei225 futures contracts. It then uses the conditional distribution for these contracts to estimate spectral risk measures, which are coherent risk measures …
Estimating properties of discrete distributions is a fundamental problem in statistical learning. We design the first unified, linear-time, competitive, property estimator that for a wide class of properties and for all underlying distributions uses just samples to achieve the performance attained by the empirical…
This work presents the concept of kernel mean embedding and kernel probabilistic programming in the context of stochastic systems. We propose formulations to represent, compare, and propagate uncertainties for fairly general stochastic dynamics in a distribution-free manner. The new tools enjoy sound theory rooted in f…
Transformer improves parameter estimation without needing closed-form solutions.
Study compares two methods for predicting extreme atmospheric events.
New method for distributed online learning with communication constraints reduces joint regret.
The Bouncy Particle Sampler is a novel rejection-free non-reversible sampler for differentiable probability distributions over continuous variables. We generalize the algorithm to piecewise differentiable distributions and apply it to generic binary distributions using a piecewise differentiable augmentation. We illust…
A new method predicts precipitation distributions from ensemble forecasts.