Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

138276414552 · Jun 202019922001200920172026
48 results for tail sampling

This work extends diffusion models to handle heavy-tailed targets, improving score estimation and sampling guarantees.

problem Score estimation and sampling guarantees for heavy-tailed targets in diffusion models.
method Kernel density estimation and minimax rates analysis for score estimation and sampling guarantees.
result Sharp minimax rates for score estimation and sampling guarantees for heavy-tailed targets, revealing qualitative differences between exponential and polynomial tails.

Study optimizes sampling to avoid extreme tail risks in unknown heavy-tailed distributions.

problem Identify optimal alternative with minimal extreme tail risk from unknown heavy-tailed distributions.
method Data-driven sequential sampling policies to maximize likelihood of selecting the optimal alternative.
result Proposed methods outperform existing approaches in identifying the optimal alternative.

Study on U-statistics with heavy-tailed samples, providing tail bounds and LDP.

problem Deviation of U-statistics with heavy-tailed samples.
method Exponential tail bounds and Large Deviation Principle (LDP) for U-statistics.
result Obtained an exponential upper bound for U-statistics tail decay, showing two regions of decay.

This paper compares VaR estimation methods under tail misspecification, finding importance sampling underestimates VaR.

problem Tail misspecification in VaR estimation.
method Importance sampling and moment-based VaR bracketing.
result Importance sampling underestimates VaR under heavy-tailed returns, while moment-based methods are robust.

Privacy affects how much data is needed for CVaR optimization.

problem Privacy constraints impact the effective sample size for CVaR optimization.
method Analyzes the privacy-relevant sample size and decomposes CVaR excess risk.
result The effective private tail sample size is εnτ, affecting CVaR learning rates.

Efficiently estimates sparse linear regression with heavy-tailed and outlier-contaminated data.

problem Estimating sparse linear regression coefficients with heavy-tailed and outlier-contaminated data.
method Efficient computation of estimators with sharp error bounds.
result Sharp error bounds for efficient estimators.

HTFM improves mode coverage and tail-statistic recovery for heavy-tailed data.

problem Tackles heavy-tailed data in various domains with rare events.
method Proposes a framework using clock-conditioned Gaussian sources and truncated logsignature features.
result Improves mode coverage, sample quality, and tail-statistic recovery over Gaussian flow matching and baselines.

In this paper we propose a new approach to estimation of the tail exponent in financial stock markets. We begin the study with the finite sample behavior of the Hill estimator under α-stable distributions. Using large Monte Carlo simulations, we show that the Hill estimator overestimates the true tail exponent and can …

2012-01-23abs ↗pdf ↗

AIS algorithm improves heavy-tailed distribution estimation.

problem Inconsistent estimators and slow convergence in AIS for heavy-tailed distributions.
method Adapts Student-t proposal distributions by matching escort moments and minimizing α-divergence.
result Improves estimation accuracy for heavy-tailed distributions.

New study shows Gaussian samplers struggle with heavy-tailed targets, while stable samplers excel.

problem The difficulty of sampling from heavy-tailed distributions using Gaussian versus stable oracles.
method Comparison of Gaussian and stable oracles for proximal samplers.
result Gaussian samplers have a fundamental barrier for high-accuracy guarantees in heavy-tailed sampling, while stable samplers excel.

In risk management, tail risks are of crucial importance. The assessment of risks should be carried out in accordance with the regulatory authority's requirement at high quantiles. In general, the underlying distribution function is unknown, the database is sparse, and therefore special tail models are used. Very often…

2019-04-27abs ↗pdf ↗

This paper analyzes sampling from heavy-tailed distributions using discretized Itô diffusions.

problem Sampling from heavy-tailed distributions with finite variance.
method Mean-square analysis of discretized Itô diffusions with weighted Poincaré inequalities.
result Explicit iteration complexity for obtaining samples close to target distributions in Wasserstein-2 metric.

Study shows pre-trained models can handle long-tailed relations well, improving classifier performance.

problem Challenges in long-tailed relation classification due to class imbalance.
method Used instance-balanced sampling to pre-train models and then improved classifier performance through attentive relation routing.
result Robust classifier with attentive relation routing achieves better long-tailed classification ability.

This research tackles imbalanced continual learning with a new sampling strategy.

problem Long-tailed distribution in multi-label datasets.
method Partitioning Reservoir Sampling (PRS) for balanced knowledge of head and tail classes.
result The proposed PRS strategy maintains a balanced knowledge of both head and tail classes.

New framework removes harmful momentum effect for long-tailed classification.

problem Challenges in maintaining balanced datasets with long-tailed data.
method Causal inference framework to disentangle and remove harmful effects of momentum.
result Achieves state-of-the-art performance on long-tailed visual recognition benchmarks.

The paper explores tail diversification in financial markets using entropy and mutual information.

problem Tail diversification in financial time series.
method Statistical independence through differential entropy and mutual information, using moments as contrast functions.
result Tail covariance matrix is a key driver of tail diversification.

Muon optimizes Transformer training with heavy-tailed data, achieving optimal sample complexity.

problem Theoretical understanding of non-Euclidean optimisation methods for heavy-tailed data in training Transformers.
method Addressing the gap in theoretical understanding, we show Muon achieves optimal sample complexity under heavy-tailed noise.
result Muon finds an ε-stationary point in nuclear norm with optimal sample complexity, absorbing heavy-tailed noise without dimension dependence.

The hidden tail of empirical distributions is analyzed using extreme value theory.

problem Understanding the bias between in-sample mean and true statistical mean for large nn.
method Extreme value theory applied to empirical distributions and their moments.
result The hidden moment of order 0 for power law distributions follows an exponential distribution with expectation 1/n1/n.

Deep models can't generate heavy-tailed samples well.

problem Understanding the limitations of deep generative models in generating samples with heavy tails.
method Unified framework using concentration of measure and convex geometry, Gromov-Levy inequality.
result Deep generative models are not universal generators and can only produce concentrated samples with light tails.

Unified approach for sampling non-differentiable and heavy-tailed targets.

problem Sampling non-differentiable and heavy-tailed distributions using Langevin algorithms.
method Anchored Langevin dynamics, which modifies the Langevin diffusion with a smooth reference potential and multiplicative scaling.
result Non-asymptotic guarantees in the 2-Wasserstein distance to the target distribution.

New bounds on private mean estimation for heavy-tailed distributions.

problem Estimating the mean of heavy-tailed distributions under differential privacy constraints.
method Upper and lower bounds on sample complexity for differentially private mean estimation.
result Qualitatively different sample complexity compared to non-private estimation, with a factor of O(d)O(d) larger for multivariate cases.

DRAGON improves learning for rare classes in unbalanced datasets using class descriptions.

problem Learning rare classes in unbalanced datasets with deep models.
method DRAGON is a late-fusion architecture that corrects bias towards frequent classes and fuses class-descriptions to improve tail-class accuracy.
result DRAGON outperforms state-of-the-art models on new benchmarks for long-tail learning with class descriptors.

The paper establishes CLTs for Markov chains and improves sampling algorithms for heavy-tailed distributions.

problem Establishing central limit theorems for ergodic averages of Markov chains.
method Drift conditions to provide necessary and sufficient conditions for CLTs, including lower bounds on convergence rates.
result Sharp conditions and convergence rates for various MCMC algorithms on heavy-tailed targets.

The study compares VaR and ES models for tail risk of electricity futures, finding AR(1)-GARCH(1,1) with Student-t distribution best.

problem Modeling tail risk of electricity futures contracts in various markets.
method Comparison of VaR and ES models using AR(1)-GARCH(1,1) with Student-t distribution, historical simulation, and quantile regression.
result AR(1)-GARCH(1,1) with Student-t distribution is the best-performing model for tail risk estimation.

Study on heavy tails in closing auction returns, explaining imbalance through limit order submission.

problem Understanding heavy tails in closing auction return distributions.
method Used the stochastic call auction model of Derksen et al. (2020a) to derive and verify a relation between tail exponents.
result Large closing price fluctuations are not caused by large market orders, but by imbalance in limit orders.

SS-GEN simulates rare events in heavy and light-tailed data.

problem Estimating probabilities of extreme events in multivariate data.
method Self-Similar Generative Estimation (SS-GEN) decomposes tail distribution into radial and angular components.
result SS-GEN generates representative extreme scenarios and estimates rare-event probabilities beyond observed data.

Efficiently estimates sparse linear regression with heavy-tailed data and outliers.

problem Sparse estimation of linear regression coefficients with heavy-tailed covariates and noises, including outliers.
method Efficient computation of robust estimator with nearly optimal error bound.
result Nearly optimal error bound for robust sparse estimation.

New method for cross-validation in high-dimensional data with dependent or heavy-tailed covariates.

problem Inconsistent cross-validation in high-dimensional settings with dependent or heavy-tailed covariates.
method ROTI-GCV framework for cross-validation under proportional asymptotics regime.
result Demonstrated accuracy of ROTI-GCV in synthetic and semi-synthetic settings.

The paper optimizes portfolios using a new GARCH model with regime switching and tempered stable innovations.

problem Mitigating left tail risk in multi-asset portfolios.
method Proposes a Markov regime-switching GARCH model with multivariate normal tempered stable innovation (MRS-MNTS-GARCH) for portfolio optimization.
result Optimal portfolios with tail risk measures outperform standard deviation-based portfolios and equally weighted portfolios in various performance metrics.

TailGAN uses GANs to detect anomalies near data distribution tails.

problem Anomaly detection near data distribution tails with current GAN limitations.
method TailGAN leverages GANs with maximum entropy regularization to generate and detect anomalies near data distribution tails.
result TailGAN achieves competitive performance on various datasets compared to existing methods.