Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

6.9%13.9%20.8%27.8% · Feb 202519922001200920182026
48 results for probability analysis

Paper simplifies calculating causation probabilities and ranks root causes.

problem Computational challenges in assessing causal relationships.
method Algorithmic simplifications and novel methodological framework for Root Cause Analysis.
result Significantly reduces computational complexity for calculating causation probabilities.

Study on optimal rates for sequential probability assignment using smoothed analysis.

problem Optimal rates for sequential probability assignment under smoothed adversaries.
method General-purpose reduction from minimax rates to transductive learning, development of an efficient algorithm using MLE oracle.
result Optimal (logarithmic) fast rates for parametric and finite VC dimension classes, sublinear regret for general classes.

Paper proposes a deep learning method to estimate fill probabilities of limit orders in LOBs.

problem Estimating the fill probabilities of limit orders in different levels of a limit order book.
method Survival analysis model using a convolutional-Transformer encoder and a monotonic neural network decoder.
result The proposed method significantly outperforms other approaches in survival analysis.

This paper develops GPCA for probability distributions using Otto-Wasserstein geometry.

problem Analyzing modes of variation in datasets of probability measures.
method Geodesic Principal Component Analysis (GPCA) on Wasserstein space with neural networks.
result Identification of geodesic curves that capture modes of variation in probability distributions.

Deep Recurrent Survival Analysis models for better event prediction and survival rate estimation.

problem Survival analysis challenges in handling data censorship and sequential patterns.
method Combines deep learning for conditional probability prediction and survival analysis for censorship handling.
result Significantly outperforms state-of-the-art solutions in various metrics on real-world tasks.

Paper analyzes convergence of ODE samplers in Wasserstein distances.

problem Limited theoretical understanding of convergence properties of probability flow ODEs.
method Convergence analysis for general probability flow ODEs in 2-Wasserstein distance.
result First non-asymptotic convergence analysis for probability flow ODE samplers.

In this paper we present a novel approach for firm default probability estimation. The methodology is based on multivariate contingent claim analysis and pair copula constructions. For each considered firm, balance sheet data are used to assess the asset value, and to compute its default probability. The asset pricing …

2014-05-06abs ↗pdf ↗

Method quantifies sensitivity of reliability analysis to uncertainty sources.

problem Computational expense in reliability analysis of complex models.
method Gaussian process surrogate model, active learning, sensitivity analysis.
result Reduces main source of error in estimating rare event probabilities.

Study on stock returns tail probabilities using stochastic volatility models.

problem Understanding tail probabilities of stock returns in stochastic volatility models.
method Analyzes stochastic differential equations for volatility, applies dimensional analysis, and uses Kolmogorov forward equation.
result Tail probabilities for short-term returns fall off like an inverse cubic and scale with the measurement interval to the power 3/2.

Bayesian approach approximates probability functions of Gaussian mixtures.

problem Approximating probability functions of non-spherical Gaussian mixtures.
method Bayesian decomposition, spherical radial decomposition, random sampling.
result Established differentiability and integral representation of gradient for probability functions.

Develops a new model for synthesizing and analyzing probability measures.

problem Synthesis and analysis of probability measures.
method Linear barycentric coding model (LBCM) using linear optimal transport (LOT) metric.
result Closed-form solution to 2-Wasserstein barycenters for compatible measures.

Paper introduces symmetric divergence link models for probability distributions.

problem Symmetric divergence measures for probability distributions.
method Two general classes of link models: one for survival functions and another for cumulative probability distribution functions.
result Advantages of symmetric divergence measures over asymmetric measures for model averaging and feature assessment.

Paper examines stability of Bayesian posterior measures using integral probability metrics.

problem Stability of Bayesian inference in large-scale inverse problems.
method New families of integral probability metrics for likelihood and prior perturbations.
result Constructs new stability results for Bayesian posterior measures.

The Wasserstein metric is an important measure of distance between probability distributions, with applications in machine learning, statistics, probability theory, and data analysis. This paper provides upper and lower bounds on statistical minimax rates for the problem of estimating a probability distribution under W…

2018-02-24abs ↗pdf ↗

The problem of categorical data analysis in high dimensions is considered. A discussion of the fundamental difficulties of probability modeling is provided, and a solution to the derivation of high dimensional probability distributions based on Bayesian learning of clique tree decomposition is presented. The main contr…

2017-08-23abs ↗pdf ↗

Develops asymptotic analysis for RandNLA sampling estimators in least-squares problems.

problem Lack of distributional information for RandNLA estimators in statistical inference.
method Asymptotic analysis of sampling estimators for least-squares problems in two settings.
result Sampling estimators are asymptotically normally distributed under mild conditions.

This work extends stochastic localization to joint probability measures for data analysis.

problem Data distributional analysis in high-dimensional probability.
method Unified stochastic localization under Eldan's α-scheme, coupled probability measures via shared Brownian motion.
result Eldan's α-distance as a scalable surrogate for Wasserstein distance.

The paper analyzes how behavioral investors make portfolio decisions using Markowitz Stochastic Dominance criteria.

problem Understanding how behavioral investors make portfolio decisions.
method Developed stochastic optimization problems and MILP models to capture subjective decision weights and probability weighting functions.
result The developed models can be used to formulate computationally tractable portfolio analysis problems.

CNFs learn distributions from samples with error bounds.

problem Learning probability distributions from finite samples.
method Continuous normalizing flows with linear interpolation and flow matching objective function.
result Non-asymptotic error bounds for distribution estimator in Wasserstein-2 distance.

Research reveals simplicity bias in random logistic map, impacting data analysis and forecasting.

problem Simplicity bias in dynamical systems and its impact on data analysis and prediction.
method Examined the logistic map and random logistic map, focusing on simplicity bias and noise effects.
result Simplicity bias is observable in the random logistic map, persisting even with small noise levels.

New insights on active sequential prediction for mean estimation.

problem Active sequential prediction-powered mean estimation problem.
method Combining uncertainty-based suggestion with a constant probability, analyzing non-asymptotic bounds, and using no-regret learning.
result The optimal query probability is close to the constraint when using no-regret learning.

New complexity analysis for estimating normalizing constants in high dimensions.

problem Estimating the normalizing constant of unnormalized probability densities in high dimensions.
method Analyze and derive the oracle complexity of annealed importance sampling.
result Oracle complexity of $\widetilde{O}\left(\frac{dβ^2{\mathcal{A}}^2}{\varepsilon^4} ight)$ for estimating ZZ within ε\varepsilon relative error.

Uncertainty-aware PCA preserves data uncertainty during dimensionality reduction.

problem Uncertainty in data affects traditional PCA methods, leading to inaccurate results.
method Generalizes PCA for multivariate probability distributions, respecting uncertainty.
result Uncertainty-aware PCA maintains data characteristics after projection.

OPAA estimates probability densities using functional analysis.

problem Estimating probability density functions efficiently and accurately.
method OPAA uses a parallelizable algorithm based on functional analysis to estimate probability distributions.
result OPAA provides an efficient method to estimate probability density functions and normalizing weights.

Probability versions of Li-Yau inequalities for manifolds with boundary.

problem Establishing Li-Yau inequalities for manifolds with non-convex boundaries.
method Stochastic analysis and Bakry-Emery curvature-dimension approach.
result Explicit probability versions of Li-Yau inequalities for manifolds with boundary.

New formula classifies product reviews into higher and lower ratings based on sentiment analysis.

problem Lack of research on using sentiment analysis for classifying text into ratings.
method Redefined sentiment proportions as a triangle structure to derive variables for classifying text into higher and lower ratings.
result Proved a dependence exists between sentiments and ratings.

This study analyzes how well GANs approximate distributions from small samples.

problem Understanding how well GANs approximate distributions from limited data.
method Analysis of GANs using integral probability metrics and Hölder classes.
result GANs can adaptively learn low-dimensional structures or Hölder densities.

ES reduces high-probability regret in stochastic linear bandits.

problem High-probability regret in stochastic linear bandits.
method Linear ensemble sampling with standard Gaussian perturbations, analyzing m=Θ(dlogn)m=Θ(d\log n) ensemble size.
result ES achieves ildeO(d3/2n) ilde O(d^{3/2}\sqrt n) high-probability regret, closing the gap to Thompson sampling.

We present a relatively detailed analysis of the persistence probability distributions in financial dynamics. Compared with the auto-correlation function, the persistence probability distributions describe dynamic correlations non-local in time. Universal and non-universal behaviors of the German DAX and Shanghai Index…

2005-11-23abs ↗pdf ↗

Paper proposes an efficient AL-GP method for CDF/CCDF estimation in UQ.

problem Estimating full probability distribution in forward UQ analysis.
method Active learning-based Gaussian process (AL-GP) metamodelling method.
result Efficient estimation of CDF/CCDF without explicit discretization.

New analysis shows halting time is predictable for large models, improving optimization efficiency.

problem Understanding the average-case complexity of optimization algorithms for large-scale models.
method Average-case analysis of first-order methods on random least squares and neural networks.
result Halting time is independent of input distribution, leading to tighter convergence rates.

The statistical properties of a stochastic process may be described (1)by the expectation values of the observables, (2)by the probability distribution functions or (3)by probability measures on path space. Here an analysis of level (3) is carried out for market fluctuation processes. Gibbs measures and chains with com…

2001-02-16abs ↗pdf ↗

Study optimizes tree-based models for better alignment of predicted scores and actual probabilities.

problem Traditional calibration metrics fail to align predicted scores with actual probabilities when score distributions deviate from the underlying data.
method Optimizes tree-based models (Random Forest, XGBoost) using Kullback-Leibler (KL) divergence to minimize the difference between predicted and true probability distributions.
result Optimized tree-based models yield superior alignment between predicted scores and actual probabilities without significant performance loss.

Entropy-based GP adaptive design improves failure probability estimation.

problem Limited accuracy in failure probability estimation due to model evaluation costs.
method Entropy-based Gaussian process (GP) adaptive design combined with multifidelity importance sampling (MFIS).
result More accurate failure probability estimates and higher confidence.