Paper simplifies calculating causation probabilities and ranks root causes.
problem Computational challenges in assessing causal relationships.
method Algorithmic simplifications and novel methodological framework for Root Cause Analysis.
result Significantly reduces computational complexity for calculating causation probabilities.
Study on the probability of immunity and its bounds.
problem Estimating the probability of immunity and its bounds.
method Derive necessary and sufficient conditions for non-immunity and ε-bounded immunity; introduce indirect immunity; propose sensitivity analysis.
result Estimate the probability of benefit and produce tighter bounds of the probability of benefit.
Meta-analysis finds people value insurance for low-probability risks more than expected.
problem Low-probability risks insurance demand is lower than expected.
method Conducted a meta-analysis of contingent valuation studies.
result Average stated willingness to pay (WTP) for insurance is 87% of expected losses.
Study on optimal rates for sequential probability assignment using smoothed analysis.
problem Optimal rates for sequential probability assignment under smoothed adversaries.
method General-purpose reduction from minimax rates to transductive learning, development of an efficient algorithm using MLE oracle.
result Optimal (logarithmic) fast rates for parametric and finite VC dimension classes, sublinear regret for general classes.
Paper proposes a deep learning method to estimate fill probabilities of limit orders in LOBs.
problem Estimating the fill probabilities of limit orders in different levels of a limit order book.
method Survival analysis model using a convolutional-Transformer encoder and a monotonic neural network decoder.
result The proposed method significantly outperforms other approaches in survival analysis.
This paper develops GPCA for probability distributions using Otto-Wasserstein geometry.
problem Analyzing modes of variation in datasets of probability measures.
method Geodesic Principal Component Analysis (GPCA) on Wasserstein space with neural networks.
result Identification of geodesic curves that capture modes of variation in probability distributions.
Deep Recurrent Survival Analysis models for better event prediction and survival rate estimation.
problem Survival analysis challenges in handling data censorship and sequential patterns.
method Combines deep learning for conditional probability prediction and survival analysis for censorship handling.
result Significantly outperforms state-of-the-art solutions in various metrics on real-world tasks.
Paper analyzes convergence of ODE samplers in Wasserstein distances.
problem Limited theoretical understanding of convergence properties of probability flow ODEs.
method Convergence analysis for general probability flow ODEs in 2-Wasserstein distance.
result First non-asymptotic convergence analysis for probability flow ODE samplers.
In this paper we present a novel approach for firm default probability estimation. The methodology is based on multivariate contingent claim analysis and pair copula constructions. For each considered firm, balance sheet data are used to assess the asset value, and to compute its default probability. The asset pricing …
Method quantifies sensitivity of reliability analysis to uncertainty sources.
problem Computational expense in reliability analysis of complex models.
method Gaussian process surrogate model, active learning, sensitivity analysis.
result Reduces main source of error in estimating rare event probabilities.
New Fourier analysis method for non-uniform Boolean hypercube.
problem Non-uniform probability measures on the Boolean hypercube.
method ANOVA-based decomposition, explicit basis, least squares problem.
result Generalization of Fourier analysis for arbitrary probability measures.
Identifies most probable flows for Kunita SDEs in fluid dynamics.
problem Modeling stochastic processes with Eulerian noise and deterministic drifts.
method Equipping the domain with a Riemannian metric from the noise, solving the resulting PDEs.
result Most probable flows differ from deterministic flows, especially under noise.
Study on stock returns tail probabilities using stochastic volatility models.
problem Understanding tail probabilities of stock returns in stochastic volatility models.
method Analyzes stochastic differential equations for volatility, applies dimensional analysis, and uses Kolmogorov forward equation.
result Tail probabilities for short-term returns fall off like an inverse cubic and scale with the measurement interval to the power 3/2.
Bayesian approach approximates probability functions of Gaussian mixtures.
problem Approximating probability functions of non-spherical Gaussian mixtures.
method Bayesian decomposition, spherical radial decomposition, random sampling.
result Established differentiability and integral representation of gradient for probability functions.
Develops a new model for synthesizing and analyzing probability measures.
problem Synthesis and analysis of probability measures.
method Linear barycentric coding model (LBCM) using linear optimal transport (LOT) metric.
result Closed-form solution to 2-Wasserstein barycenters for compatible measures.
Paper introduces symmetric divergence link models for probability distributions.
problem Symmetric divergence measures for probability distributions.
method Two general classes of link models: one for survival functions and another for cumulative probability distribution functions.
result Advantages of symmetric divergence measures over asymmetric measures for model averaging and feature assessment.
Paper examines stability of Bayesian posterior measures using integral probability metrics.
problem Stability of Bayesian inference in large-scale inverse problems.
method New families of integral probability metrics for likelihood and prior perturbations.
result Constructs new stability results for Bayesian posterior measures.
Study classifies submanifolds in probability simplex.
problem Classifying submanifolds in the probability simplex.
method Complete classification through geometric analysis.
result Doubly totally-umbilical submanifolds identified and classified.
The Wasserstein metric is an important measure of distance between probability distributions, with applications in machine learning, statistics, probability theory, and data analysis. This paper provides upper and lower bounds on statistical minimax rates for the problem of estimating a probability distribution under W…
The problem of categorical data analysis in high dimensions is considered. A discussion of the fundamental difficulties of probability modeling is provided, and a solution to the derivation of high dimensional probability distributions based on Bayesian learning of clique tree decomposition is presented. The main contr…
Paper analyzes high probability convergence of adaptive SGD with momentum.
problem Theoretical understanding of adaptive SGD with momentum in nonconvex settings is incomplete.
method High probability analysis under weak assumptions.
result First high probability convergence proof for gradients to zero in Delayed AdaGrad with momentum.
Refined analysis of Mitra's algorithm for discrete mixtures.
problem Classifying general discrete mixture distribution models.
method Spectral clustering tailored to bipartite stochastic block models.
result Improved separation conditions for probability distributions.
Formally proves machine learning for simple classifiers.
problem Proving PAC learnability for decision stumps.
method Formal proof in Lean, separating deterministic and probabilistic proofs.
result Formal proof of PAC learnability for decision stumps.
Develops asymptotic analysis for RandNLA sampling estimators in least-squares problems.
problem Lack of distributional information for RandNLA estimators in statistical inference.
method Asymptotic analysis of sampling estimators for least-squares problems in two settings.
result Sampling estimators are asymptotically normally distributed under mild conditions.
This work extends stochastic localization to joint probability measures for data analysis.
problem Data distributional analysis in high-dimensional probability.
method Unified stochastic localization under Eldan's α-scheme, coupled probability measures via shared Brownian motion.
result Eldan's α-distance as a scalable surrogate for Wasserstein distance.
The paper analyzes how behavioral investors make portfolio decisions using Markowitz Stochastic Dominance criteria.
problem Understanding how behavioral investors make portfolio decisions.
method Developed stochastic optimization problems and MILP models to capture subjective decision weights and probability weighting functions.
result The developed models can be used to formulate computationally tractable portfolio analysis problems.
CNFs learn distributions from samples with error bounds.
problem Learning probability distributions from finite samples.
method Continuous normalizing flows with linear interpolation and flow matching objective function.
result Non-asymptotic error bounds for distribution estimator in Wasserstein-2 distance.
Research reveals simplicity bias in random logistic map, impacting data analysis and forecasting.
problem Simplicity bias in dynamical systems and its impact on data analysis and prediction.
method Examined the logistic map and random logistic map, focusing on simplicity bias and noise effects.
result Simplicity bias is observable in the random logistic map, persisting even with small noise levels.
New insights on active sequential prediction for mean estimation.
problem Active sequential prediction-powered mean estimation problem.
method Combining uncertainty-based suggestion with a constant probability, analyzing non-asymptotic bounds, and using no-regret learning.
result The optimal query probability is close to the constraint when using no-regret learning.
This review explores entropy applications in data analysis and machine learning.
problem Characterizing probability mass distributions in data analysis and machine learning.
method Review of various entropy types and their applications.
result Entropy's versatility in data analysis and machine learning.
New complexity analysis for estimating normalizing constants in high dimensions.
problem Estimating the normalizing constant of unnormalized probability densities in high dimensions.
method Analyze and derive the oracle complexity of annealed importance sampling.
result Oracle complexity of $\widetilde{O}\left(\frac{dβ^2{\mathcal{A}}^2}{\varepsilon^4}
ight)$ for estimating Z Z Z within ε \varepsilon ε relative error. Uncertainty-aware PCA preserves data uncertainty during dimensionality reduction.
problem Uncertainty in data affects traditional PCA methods, leading to inaccurate results.
method Generalizes PCA for multivariate probability distributions, respecting uncertainty.
result Uncertainty-aware PCA maintains data characteristics after projection.
OPAA estimates probability densities using functional analysis.
problem Estimating probability density functions efficiently and accurately.
method OPAA uses a parallelizable algorithm based on functional analysis to estimate probability distributions.
result OPAA provides an efficient method to estimate probability density functions and normalizing weights.
There are many advantages to use probability method for nonlinear system identification, such as the noises and outliers in the data set do not affect the probability models significantly; the input features can be extracted in probability forms. The biggest obstacle of the probability model is the probability distribu…
Probability versions of Li-Yau inequalities for manifolds with boundary.
problem Establishing Li-Yau inequalities for manifolds with non-convex boundaries.
method Stochastic analysis and Bakry-Emery curvature-dimension approach.
result Explicit probability versions of Li-Yau inequalities for manifolds with boundary.
New formula classifies product reviews into higher and lower ratings based on sentiment analysis.
problem Lack of research on using sentiment analysis for classifying text into ratings.
method Redefined sentiment proportions as a triangle structure to derive variables for classifying text into higher and lower ratings.
result Proved a dependence exists between sentiments and ratings.
We prove a general theorem providing smoothed analysis estimates for conic condition numbers of problems of numerical analysis. Our probability estimates depend only on geometric invariants of the corresponding sets of ill-posed inputs. Several applications to linear and polynomial equation solving show that the estima…
This study analyzes how well GANs approximate distributions from small samples.
problem Understanding how well GANs approximate distributions from limited data.
method Analysis of GANs using integral probability metrics and Hölder classes.
result GANs can adaptively learn low-dimensional structures or Hölder densities.
ES reduces high-probability regret in stochastic linear bandits.
problem High-probability regret in stochastic linear bandits.
method Linear ensemble sampling with standard Gaussian perturbations, analyzing m = Θ ( d log n ) m=Θ(d\log n) m = Θ ( d log n ) ensemble size. result ES achieves i l d e O ( d 3 / 2 n ) ilde O(d^{3/2}\sqrt n) i l d e O ( d 3/2 n ) high-probability regret, closing the gap to Thompson sampling. We present a relatively detailed analysis of the persistence probability distributions in financial dynamics. Compared with the auto-correlation function, the persistence probability distributions describe dynamic correlations non-local in time. Universal and non-universal behaviors of the German DAX and Shanghai Index…
Categorical d-separation criterion simplifies probability graph analysis.
problem Detecting causal relationships in probability distributions.
method Introducing categorical definitions for causal models and d-separation.
result Abstract version of d-separation criterion applies to various probability theories.
Paper proposes an efficient AL-GP method for CDF/CCDF estimation in UQ.
problem Estimating full probability distribution in forward UQ analysis.
method Active learning-based Gaussian process (AL-GP) metamodelling method.
result Efficient estimation of CDF/CCDF without explicit discretization.
PGF kernels analyze spherical data using generalized RBF kernels.
problem Analysis of spherical data.
method Introduced PGF kernels and a semi-parametric learning algorithm.
result PGF kernels generalize RBF kernels for spherical data.
New analysis shows halting time is predictable for large models, improving optimization efficiency.
problem Understanding the average-case complexity of optimization algorithms for large-scale models.
method Average-case analysis of first-order methods on random least squares and neural networks.
result Halting time is independent of input distribution, leading to tighter convergence rates.
The statistical properties of a stochastic process may be described (1)by the expectation values of the observables, (2)by the probability distribution functions or (3)by probability measures on path space. Here an analysis of level (3) is carried out for market fluctuation processes. Gibbs measures and chains with com…
We give the proof of a tight lower bound on the probability that a binomial random variable exceeds its expected value. The inequality plays an important role in a variety of contexts, including the analysis of relative deviation bounds in learning theory and generalization bounds for unbounded loss functions.
Study optimizes tree-based models for better alignment of predicted scores and actual probabilities.
problem Traditional calibration metrics fail to align predicted scores with actual probabilities when score distributions deviate from the underlying data.
method Optimizes tree-based models (Random Forest, XGBoost) using Kullback-Leibler (KL) divergence to minimize the difference between predicted and true probability distributions.
result Optimized tree-based models yield superior alignment between predicted scores and actual probabilities without significant performance loss.
Entropy-based GP adaptive design improves failure probability estimation.
problem Limited accuracy in failure probability estimation due to model evaluation costs.
method Entropy-based Gaussian process (GP) adaptive design combined with multifidelity importance sampling (MFIS).
result More accurate failure probability estimates and higher confidence.