Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

207415622829 · Jun 202019922001200920172026
48 results for Empirical Distribution

Neural Empirical Bayes estimates source distributions from noisy simulations.

problem Estimating source distributions from noisy, simulated data.
method Uses neural density estimators to estimate a prior or source distribution over uncorrupted samples, then performs posterior inference.
result Recovering ground truth source distributions up to symmetries.

Personal income distributions in Japan are analyzed empirically and a simple stochastic model of the income process is proposed. Based on empirical facts, we propose a minimal two-factor model. Our model of personal income consists of an asset accumulation process and a wage process. We show that these simple processes…

2005-05-25abs ↗pdf ↗

Improved bounds for discrete probability distribution estimation under the ℓ∞ norm.

problem Estimating discrete probability distributions under the ℓ∞ norm with improved bounds.
method Minimax bounds in expectation and high-probability tail bounds.
result Resolved open questions posed in Kontorovich and Painsky (JMLR, 2025), including a fully empirical tightest risk bound and identifying the worst-case extremal distribution.

Exact distribution of split conformal prediction coverage found.

problem Determining the reliability of prediction sets in batch mode.
method Analysis of exchangeable data to find universal distribution of empirical coverage.
result Exact distribution of empirical coverage is universal and determined by nominal miscoverage level and calibration sample size.

We study 'meta-dependence' in conditional independence tests across different empirical distributions.

problem Understanding the breakdown of conditional independence properties in finite data.
method Geometric intuition and information projections to measure meta-dependence between conditional independences.
result We provide a measure of meta-dependence that consolidates findings across synthetic and real-world data.

Sharp bounds for max-sliced Wasserstein distances derived for empirical distributions.

problem Estimating the expected max-sliced Wasserstein distance between a probability measure and its empirical distribution.
method Banach space version and operator norm approach for upper bounds.
result Upper bounds for max-sliced Wasserstein distances are essentially matching and sharp up to a log factor.

ECOD detects outliers without parameters, fast and simple.

problem Detecting outliers in large, high-dimensional datasets efficiently and interpretably.
method ECOD estimates empirical cumulative distribution functions per dimension, then computes tail probabilities and outlier scores.
result ECOD outperforms state-of-the-art methods in accuracy, efficiency, and scalability.

We consider a financial market model which consists of a financial asset and a large number of interacting agents classified into many types. Different types of agents are heterogeneous in their price expectations. Each agent can change its type based on the current empirical distribution of the types and the equilibri…

2007-03-28abs ↗pdf ↗

The best-known and most commonly used distribution-property estimation technique uses a plug-in estimator, with empirical frequency replacing the underlying distribution. We present novel linear-time-computable estimators that significantly "amplify" the effective amount of data available. For a large variety of distri…

2019-03-04abs ↗pdf ↗

Enhances flexibility in data reweighting with optimal transport and maximum entropy principles.

problem Adapting empirical distributions to predefined constraints on moments, tail behavior, etc.
method Nonparametric distributional constraints, maximum entropy principle, optimal transport.
result Maximum entropy weight adjusted empirical distribution close to a specified distribution in optimal transport metric.

This work studies the smooth 1-Wasserstein distance and its limit distribution in high dimensions.

problem Addressing the curse of dimensionality in empirical approximation.
method Conducts a statistical study including limit distribution, bootstrap consistency, and concentration inequalities.
result Derives a nondegenerate limit distribution for empirical SWD, contrasting with classic W1W_1.

The paper improves generative models to avoid replicating observed examples.

problem Improving generative models to avoid replicating observed examples.
method Theoretical insights into the Wasserstein GAN, constrained to left-invertible push-forward maps, generating distributions that avoid replication and significantly deviate from the empirical distribution.
result Left-invertibility achieves this without compromising statistical optimality.

The hidden tail of empirical distributions is analyzed using extreme value theory.

problem Understanding the bias between in-sample mean and true statistical mean for large nn.
method Extreme value theory applied to empirical distributions and their moments.
result The hidden moment of order 0 for power law distributions follows an exponential distribution with expectation 1/n1/n.

New approach avoids excess empirical risk in domain generalization.

problem Learning models that generalize to unseen distributions from diverse data sets.
method Minimizes penalty under constraint of optimal empirical risk, leveraging rate-distortion theory.
result Significant improvements in domain generalization performance across multiple methods.

Improved eigenvalue distribution method for financial data.

problem Noise and complexity in financial markets.
method Matrix H theory, hierarchical structure, informational cascade.
result Captures a larger fraction of data variance in financial markets.

Detect changes in noisy dynamical systems using empirical approximations and finite-sample bounds.

problem Change detection in noisy dynamical systems
method Partition-based empirical approximations and finite-state stationary distribution stability
result Finite-sample bound for empirical stationary density

This paper analyzes quantiles of heavy-tailed distributions, separating projection direction and quantile threshold effects.

problem Analyzing quantiles of heavy-tailed distributions with estimated parameters.
method Introduces a Q-Q orthogonality formulation to separate projection-direction and quantile-threshold effects.
result Decomposes the difference between empirical and population quantiles into three terms.

Pareto's 80/20 rule follows a Gaussian distribution with twice the mean standard deviation.

problem Understanding variations in the 80/20 rule across different contexts.
method Identifying the statistical distribution of the 80/20 rule and its variations.
result The 80/20 rule follows a Gaussian distribution with a standard deviation twice the mean.

Develops robust MDPs for unknown disturbances with performance guarantees.

problem Unknown disturbance distribution in MDPs.
method Empirical distribution, sublevel set of distance function, weak convergence, concentration inequality.
result Robust optimal value function converges to true optimal value function with increasing sample sizes.

Distributional reinforcement learning (distributional RL) has seen empirical success in complex Markov Decision Processes (MDPs) in the setting of nonlinear function approximation. However, there are many different ways in which one can leverage the distributional approach to reinforcement learning. In this paper, we p…

2018-05-13abs ↗pdf ↗

This study examines biases in flow matching samplers using finite-sample estimation.

problem Biases in flow matching samplers when using finite-sample surrogates.
method Finite-sample plug-in estimation and hierarchy of empirical FM models.
result Exact empirical minimizer and smoothed plug-in regime identified for affine conditional flows.

Study analyzes stock market correlations using multivariate distributions.

problem Capturing the correlation structure of complex, non-stationary systems.
method Applied Random Matrix Model to empirical data of 479 US stocks.
result Described and quantified changes in empirical distributions due to non-stationarity.

Transformers can learn Markov processes with constant depth, surprising results.

problem Understanding how transformers learn context in Markov processes.
method Empirical study and theoretical analysis of attention-based transformers on Markov data.
result Transformers with constant depth can achieve low test loss on Markov sequences, matching empirical and theoretical findings.

Auto-decoder synthesizes graphs from latent codes.

problem Creating new graph structures from specified distributions.
method Generative model learns latent codes from empirical distribution. Self-attention identifies likely connectivity patterns. Graph-based normalizing flows sample latent codes.
result Model outperforms state of the art by 1.5x in accuracy and 2x in speed.

cCorrGAN approximates conditional correlation matrices using GANs.

problem Learning empirical conditional distributions in the elliptope of correlation matrices.
method Conditional Generative Adversarial Networks (GANs) applied to correlation matrices.
result Validated through Monte Carlo simulations in finance.

Transformer pretraining yields strong EB performance without explicit adaptation.

problem Empirical Bayes problems with unknown test distributions.
method Indirect analysis of pretrained transformer's performance under universal priors.
result Near-optimal regret bound of O~(1n)\widetilde{O}(\frac{1}{n}) for arbitrary test distributions.

Unsupervised domain adaptation is a promising way to generalize deep models to novel domains. However, the current literature assumes that the label distribution is domain-invariant and only aligns the feature distributions or vice versa. In this work, we explore the more realistic task of Class-imbalanced Domain Adapt…

2019-10-23abs ↗pdf ↗

We study the risk performance of distributed learning for the regularization empirical risk minimization with fast convergence rate, substantially improving the error analysis of the existing divide-and-conquer based distributed learning. An interesting theoretical finding is that the larger the diversity of each local…

2018-12-19abs ↗pdf ↗

New statistical test for change-point detection using relative entropy.

problem Offline change-point detection using divergence metrics.
method Study of empirical relative entropy distributions, derivation of approximations, introduction of new Berry-Esseen bounds.
result Theoretical and practical validation of relative entropy for change-point detection.

EB-PCA reduces noise in high-dimensional PCA by estimating a joint prior distribution.

problem High-dimensional PCA noise in samples comparable to or larger than data.
method Empirical Bayes PCA using Kiefer-Wolfowitz MLE, random matrix theory, and AMP algorithm.
result EB-PCA achieves Bayes-optimal accuracy in spiked models and significantly improves over PCA in simulations and real data.

Adaptive model learns from time series data with changing distributions.

problem Predicting time series data under distribution shift.
method Formulates distribution shift as weighted empirical risk minimization. Uses a gradient-based learning method for a forgetting mechanism.
result Proposes an efficient method for adaptive time series prediction.

The paper improves the empirical bootstrap method for non-normal estimators.

problem Theoretical properties of empirical bootstrap for non-asymptotically normal estimators.
method Establishing limiting distribution, deriving consistency conditions, proposing alternative methods.
result The empirical bootstrap method can be asymptotically consistent under stability conditions.