Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

3136279401,253 · Jun 202019922001200920172026
48 results for empirical distribution function

Develops robust MDPs for unknown disturbances with performance guarantees.

problem Unknown disturbance distribution in MDPs.
method Empirical distribution, sublevel set of distance function, weak convergence, concentration inequality.
result Robust optimal value function converges to true optimal value function with increasing sample sizes.

New method corrects bias in datasets using cumulative distribution functions.

problem Varying domains and biased datasets lead to differences between training and target distributions.
method Empirical cumulative distribution function estimates of the target distribution, rigorously generalized.
result Method is more robust, not reliant on parameter tuning, and performs similarly to state-of-the-art techniques.

Neural Empirical Bayes estimates source distributions from noisy simulations.

problem Estimating source distributions from noisy, simulated data.
method Uses neural density estimators to estimate a prior or source distribution over uncorrupted samples, then performs posterior inference.
result Recovering ground truth source distributions up to symmetries.

Study risk bounds for distributed ERM with general loss functions and hypothesis spaces.

problem Limited theoretical analysis for distributed ERM with general loss functions and hypothesis spaces.
method Derive tight risk bounds under assumptions on hypothesis space and loss function.
result Developed more general risk bound for distributed ERM without strong convexity restriction.

This paper presents a distance-based discriminative framework for learning with probability distributions. Instead of using kernel mean embeddings or generalized radial basis kernels, we introduce embeddings based on dissimilarity of distributions to some reference distributions denoted as templates. Our framework exte…

2018-03-01abs ↗pdf ↗

The paper studies empirical processes from nearest neighbors in regression.

problem Estimating conditional cumulative distribution functions and local linear regression.
method Uniform central limit theorem and non-asymptotic bound under local bracketing entropy and uniform entropy numbers.
result Gaussian limit of empirical process with simple covariance.

PVI improves SIVI by directly optimizing ELBO without parametric assumptions.

problem Intractable variational densities in SIVI methods.
method Particle Variational Inference (PVI) using empirical measures to approximate optimal mixing distributions.
result PVI directly optimizes the ELBO and performs favorably compared to other SIVI methods.

Distributional reinforcement learning (distributional RL) has seen empirical success in complex Markov Decision Processes (MDPs) in the setting of nonlinear function approximation. However, there are many different ways in which one can leverage the distributional approach to reinforcement learning. In this paper, we p…

2018-05-13abs ↗pdf ↗

The goal of regression and classification methods in supervised learning is to minimize the empirical risk, that is, the expectation of some loss function quantifying the prediction error under the empirical distribution. When facing scarce training data, overfitting is typically mitigated by adding regularization term…

2017-10-27abs ↗pdf ↗

Changes (returns) in stock index prices and exchange rates for currencies are argued, based on empirical data, to obey a stable distribution with characteristic exponent α<2 α< 2 for short sampling intervals and a Gaussian distribution for long sampling intervals. In order to explain this phenomenon, an Ehrenfest model…

2003-11-26abs ↗pdf ↗

A new model for stock price fluctuations is proposed, based upon an analogy with the motion of tracers in Gaussian random fields, as used in turbulent dispersion models and in studies of transport in dynamically disordered media. Analytical and numerical results for this model in a special limiting case of a single-sca…

2003-11-28abs ↗pdf ↗

A new method uses normalizing flows to approximate optimal transport between empirical distributions.

problem Learning an optimal transport map between two empirical distributions.
method Relaxing the Monge formulation of optimal transport, using normalizing flows to approximate the solution.
result The method provides a good approximation of the true optimal transport.

New metrics avoid high-dimensional analysis challenges, proving convergence without 'curse of dimensionality'.

problem High-dimensional analysis challenges in empirical measure convergence.
method Proposed a new class of probability metrics free of the curse of dimensionality.
result Convergence of empirical measures is free of the curse of dimensionality.

We reformulate unsupervised dimension reduction problem (UDR) in the language of tempered distributions, i.e. as a problem of approximating an empirical probability density function by another tempered distribution, supported in a kk-dimensional subspace. We show that this task is connected with another classical prob…

2019-03-12abs ↗pdf ↗

ECOD detects outliers without parameters, fast and simple.

problem Detecting outliers in large, high-dimensional datasets efficiently and interpretably.
method ECOD estimates empirical cumulative distribution functions per dimension, then computes tail probabilities and outlier scores.
result ECOD outperforms state-of-the-art methods in accuracy, efficiency, and scalability.

The paper bounds the expectation of empirical processes indexed by Hölder classes.

problem Estimating the expectation of the supremum of empirical processes for distributions on bounded sets.
method Providing upper bounds on the expectation of the supremum of empirical processes indexed by Hölder classes.
result Deriving non-asymptotic risk bounds for estimating distributions using empirical processes and IPM.

Corrects sample selection bias in empirical risk minimization using importance sampling.

problem Statistical learning with biased training data.
method Weighted empirical risk minimization using importance sampling.
result Generalization capacity preserved with estimated importance weights.

Distributed machine learning is an approach allowing different parties to learn a model over all data sets without disclosing their own data. In this paper, we propose a weighted distributed differential privacy (WD-DP) empirical risk minimization (ERM) method to train a model in distributed setting, considering differ…

2019-10-23abs ↗pdf ↗

Develops uniform convergence guarantees for a broad class of risk functionals in supervised learning.

problem Bounding generalization gaps for various risk functionals beyond the expectation.
method Establishes uniform convergence for Hölder risk functionals, providing guarantees for empirical risk minimization.
result First uniform convergence results for estimating the CDF of loss distributions, applicable to various risk functionals.

Employing profits data of Japanese companies in 2002 and 2003, we identify the non-Gibrat's law which holds in the middle profits region. From the law of detailed balance in all regions, Gibrat's law in the high region and the non-Gibrat's law in the middle region, we kinematically derive the profits distribution funct…

2005-08-24abs ↗pdf ↗

The paper examines the tilted empirical risk's generalization and robustness under negative tilt.

problem The generalization error of machine learning algorithms under negative tilt.
method Uniform and information-theoretic bounds on the tilted generalization error under negative tilt.
result The tilted empirical risk's generalization error has a convergence rate of \(O(n^{-ε/(1+ε)})\).

The paper introduces a DRM for causal inference, offering a flexible method to analyze counterfactual distributions.

problem Estimating mean causal effects is limited; a distributional perspective is needed for a more thorough understanding.
method The paper employs a semiparametric density ratio model (DRM) with an empirical likelihood (EL) approach to estimate counterfactual distribution functions.
result The DRM framework enables direct and transparent causal inference from a distributional perspective, validated by numerical studies.

Decentralized learning for GLMs with feature distribution and network connectivity.

problem Optimizing generalized linear models in a decentralized network with feature partitioning.
method Chambolle--Pock primal--dual algorithm applied to an equivalent saddle-point formulation.
result Convergence rates for empirical risk minimization under Lipschitz and square root Lipschitz assumptions.

Productions functions map the inputs of a firm or a productive system onto its outputs. This article expounds generalizations of the production function that include state variables, organizational structures and increasing returns to scale. These extensions are needed in order to explain the regularities of the empiri…

2005-11-22abs ↗pdf ↗

The study bounds the utility of empirically optimal portfolios using stock return data.

problem Maximizing expected ratio of portfolio utility to best asset utility.
method High probability utility bounds derived from Lipschitz or Hölder continuous utility functions.
result Utility bounds depend on utility function, number of assets, and observations.

Study analyzes stock market correlations using multivariate distributions.

problem Capturing the correlation structure of complex, non-stationary systems.
method Applied Random Matrix Model to empirical data of 479 US stocks.
result Described and quantified changes in empirical distributions due to non-stationarity.

In the spirit of the emergent field of econophysics, a goodness-of-fit test for the Power-Law distribution, based on the Empirical Distribution Function (EDF) is presented, and related problems are discussed. An analysis of the tail behaviour of the daily logarithmic variation of the Mexican Stock Market Index (IPC), s…

2003-03-27abs ↗pdf ↗

The paper proposes a new method for density estimation using spline quasi-interpolation for clustering.

problem Density estimation and clustering modeling for multivariate data.
method Spline quasi-interpolation for mono-variate approximation, copulas for multivariate modeling.
result The proposed method achieves accurate clustering of data using copulas and spline quasi-interpolation.

Neural networks estimate statistical divergences with performance guarantees.

problem Estimating statistical divergences with theoretical performance guarantees.
method Parametrizing empirical variational form by a neural network and optimizing over parameter space.
result Established non-asymptotic absolute error bounds for neural estimators of four f\mathsf{f}-divergences.

Uniform consistency proven for spatial distribution and depth estimators in any dimension.

problem Uniform consistency of spatial distribution and depth estimators in arbitrary dimensions.
method Proof of uniform L1L^1-consistency using sample size nn as the only dependency.
result Consistency rate is independent of dimension dd and sample size nn.

In this work we afford the statistical characterization of a linear Stochastic Volatility Model featuring Inverse Gamma stationary distribution for the instantaneous volatility. We detail the derivation of the moments of the return distribution, revealing the role of the Inverse Gamma law in the emergence of fat tails,…

2010-11-27abs ↗pdf ↗

Bayesian networks with latent variables are characterized and their likelihoods compared.

problem Characterizing and comparing likelihoods of Bayesian networks with latent variables.
method Characterized likelihood function and empirical Bayesian network. Proved dominance of global maximum likelihood from empirical model.
result The global maximum likelihood of the original Bayesian network is attained if and only if parameters are consistent with empirical model.