Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

140280419559 · Jun 202019922001200920172026
48 results for empirical probabilities

Improved bounds for discrete probability distribution estimation under the ℓ∞ norm.

problem Estimating discrete probability distributions under the ℓ∞ norm with improved bounds.
method Minimax bounds in expectation and high-probability tail bounds.
result Resolved open questions posed in Kontorovich and Painsky (JMLR, 2025), including a fully empirical tightest risk bound and identifying the worst-case extremal distribution.

Sharp bounds for max-sliced Wasserstein distances derived for empirical distributions.

problem Estimating the expected max-sliced Wasserstein distance between a probability measure and its empirical distribution.
method Banach space version and operator norm approach for upper bounds.
result Upper bounds for max-sliced Wasserstein distances are essentially matching and sharp up to a log factor.

While defaults are rare events, losses can be substantial even for credit portfolios with a large number of contracts. Therefore, not only a good evaluation of the probability of default is crucial, but also the severity of losses needs to be estimated. The recovery rate is often modeled independently with regard to th…

2012-03-14abs ↗pdf ↗

Aggregates probability models using Wasserstein space and variational approach.

problem Model aggregation in the Wasserstein space of distributions.
method Data-driven calibration framework based on ΓΓ-convergence.
result Empirical minimizers converge to the minimizers of the actual problem.

Motivated by problems of anomaly detection, this paper implements the Neyman-Pearson paradigm to deal with asymmetric errors in binary classification with a convex loss. Given a finite collection of classifiers, we combine them and obtain a new classifier that satisfies simultaneously the two following properties with …

2011-02-28abs ↗pdf ↗

This paper improves learning uncertain Bayesian networks from incomplete data.

problem Learning conditional probabilities in Bayesian networks with limited data.
method Develops methods to estimate and quantify uncertainty in conditional probabilities with incomplete data.
result Improves state-of-the-art approaches for handling uncertain Bayesian networks with incomplete data.

New method approximates high-dimensional probability densities efficiently.

problem Approximating high-dimensional probability densities accurately and efficiently.
method Hierarchical tensor-network approach using randomized SVD and linear equations.
result The method effectively approximates high-dimensional densities with linear complexity.

The paper analyzes the performance of empirical risk minimization for pp-norm linear regression.

problem Empirical risk minimization on pp-norm linear regression.
method Analyzes performance under various conditions and moment assumptions.
result High probability excess risk bounds for empirical risk minimizer, matching asymptotic rates.

We discuss the problem of risk estimation in the classification problem, with specific focus on finding distributions that maximize the confidence intervals of risk estimation. We derived simple analytic approximations for the maximum bias of empirical risk for histogram classifier. We carry out a detailed study on usi…

2014-08-14abs ↗pdf ↗

Based on empirical financial time-series, we show that the "silence-breaking" probability follows a super-universal power law: the probability of observing a large movement is inversely proportional to the length of the on-going low-variability period. Such a scaling law has been previously predicted theoretically [R. …

2008-12-24abs ↗pdf ↗

ECOD detects outliers without parameters, fast and simple.

problem Detecting outliers in large, high-dimensional datasets efficiently and interpretably.
method ECOD estimates empirical cumulative distribution functions per dimension, then computes tail probabilities and outlier scores.
result ECOD outperforms state-of-the-art methods in accuracy, efficiency, and scalability.

The paper bounds the expectation of empirical processes indexed by Hölder classes.

problem Estimating the expectation of the supremum of empirical processes for distributions on bounded sets.
method Providing upper bounds on the expectation of the supremum of empirical processes indexed by Hölder classes.
result Deriving non-asymptotic risk bounds for estimating distributions using empirical processes and IPM.

We study the minimax optimal rate for estimating the Wasserstein-11 metric between two unknown probability measures based on nn i.i.d. empirical samples from them. We show that estimating the Wasserstein metric itself between probability measures, is not significantly easier than estimating the probability measures u…

2019-08-27abs ↗pdf ↗

The paper studies PCA of probability measures with varying sample sizes and finds optimal convergence rates.

problem PCA of multiple probability measures with varying sample sizes.
method Double asymptotic regime analysis with convergence rates n1/2+mαn^{-1/2} + m^{-α} for empirical covariance and PCA risk.
result Optimal convergence rates for empirical covariance and PCA risk in the dense regime are proven.

New method for PU learning with instance-dependent propensity scores.

problem Learning from positive and unlabeled data with instance-dependent labeling.
method Empirical risk minimization of joint risk function, alternating optimization of posterior probability and propensity score.
result The method achieves comparable or better performance than state-of-the-art methods.

SURF simplifies distribution estimation with simple, robust, and fast algorithms.

problem Efficient and accurate distribution estimation in statistics and machine learning.
method Piecewise polynomial approximation using empirical probability interpolation and divide-and-conquer merging.
result Surpassing state-of-the-art algorithms in efficiency and accuracy, SURF estimates distributions robustly and quickly.

A new method for matrix completion with model-free weights.

problem Matrix completion under non-uniform missing structures.
method Constructs weights via convex optimization to adjust for non-uniformity without modeling observation probabilities.
result Recover matrix with stronger theoretical guarantees, especially in heterogeneous missing settings.

Nonparametric adaptive robust control tackles model uncertainty in stochastic processes.

problem Model uncertainty in stochastic processes.
method Adaptive robust control methodology using online learning and uncertainty reduction, empirical distribution, and Lagrangian duality.
result Nonparametric adaptive robust control approach is preferable to traditional robust frameworks.

A new dynamical formulation of log-PCA captures local principal modes of geodesic variations.

problem Learning principal variations of random probability measures under Wasserstein geometry.
method Introducing a new dynamical formulation of log-PCA as a variational approach.
result Deriving a general statistical convergence rate for empirical WT-PCA.

Investigates financial and economic systems using statistical mechanics and information theory.

problem Complexity, asymmetry, stochasticity, and non-linearity in financial and economic systems.
method Model-based and empirical analyses using statistical mechanics and information theory.
result Derives probability distribution functions for better understanding of financial and economic dynamics.

The paper proposes a new method for covariate balancing using IPM to improve causal inference.

problem Covariate imbalance in causal inference weighting methods, especially when models are not correctly specified.
method The integral probability metric (IPM) is used to determine optimal weights for treated and control groups.
result The proposed method can be consistent without specifying either the propensity score or outcome regression model.

This paper introduces AdaSDCA: an adaptive variant of stochastic dual coordinate ascent (SDCA) for solving the regularized empirical risk minimization problems. Our modification consists in allowing the method adaptively change the probability distribution over the dual variables throughout the iterative process. AdaSD…

2015-02-27abs ↗pdf ↗

New metrics avoid high-dimensional analysis challenges, proving convergence without 'curse of dimensionality'.

problem High-dimensional analysis challenges in empirical measure convergence.
method Proposed a new class of probability metrics free of the curse of dimensionality.
result Convergence of empirical measures is free of the curse of dimensionality.

We investigate the historical volatility of the 100 most capitalized stocks traded in US equity markets. An empirical probability density function (pdf) of volatility is obtained and compared with the theoretical predictions of a lognormal model and of the Hull and White model. The lognormal model well describes the pd…

2002-02-28abs ↗pdf ↗

LinMED is a new linear bandit algorithm with near-optimal regret bound.

problem Optimizing decision-making in linear bandit problems with sub-Gaussian distributions.
method LinMED is a randomized linear bandit algorithm with closed-form arm sampling probabilities.
result LinMED achieves a near-optimal regret bound of dnd\sqrt{n} up to logarithmic factors.

Boosted decision trees typically yield good accuracy, precision, and ROC area. However, because the outputs from boosting are not well calibrated posterior probabilities, boosting yields poor squared error and cross-entropy. We empirically demonstrate why AdaBoost predicts distorted probabilities and examine three cali…

2012-07-04abs ↗pdf ↗

While Gaussian probability densities are omnipresent in applied mathematics, Gaussian cumulative probabilities are hard to calculate in any but the univariate case. We study the utility of Expectation Propagation (EP) as an approximate integration method for this problem. For rectangular integration regions, the approx…

2011-11-29abs ↗pdf ↗

Asset prices contain information about the probability distribution of future states and the stochastic discounting of those states as used by investors. To better understand the challenge in distinguishing investors' beliefs from risk-adjusted discounting, we use Perron-Frobenius Theory to isolate a positive martingal…

2014-11-28abs ↗pdf ↗

This study investigates self-supervised learning with Wasserstein distance on tree structures.

problem Improving self-supervised learning methods using Wasserstein distance.
method Utilized Tree-Wasserstein distance (TWD) and Jeffrey divergence regularization for training.
result A simple combination of softmax function and Tree-Wasserstein distance outperforms cosine similarity-based methods.

New probability path model improves flow matching forecasting performance.

problem Impact of probability path model selection on flow matching forecasting performance.
method Proposed a novel probability path model designed to improve forecasting performance.
result Our model achieves faster convergence during training and improved predictive performance compared to existing models.

An algorithm finds the most probable best solution in uncertain parameter settings.

problem Finding the most probable best solution in uncertain parameter settings.
method Designing an efficient sequential sampling algorithm to learn the most probable best (MPB) and optimizing the computing budget allocation.
result The algorithms achieve the optimal sampling ratios as the simulation budget increases and significantly improve empirical performance.