Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

138276413551 · Jun 202019922001200920172026
48 results for exceedance distribution

Study compares two methods for predicting extreme atmospheric events.

problem Forecasting threshold exceedances of atmospheric variables like temperature and wind speed.
method Direct vs. full distribution probabilistic methods for rare events.
result Full distribution approach outperforms direct method for extreme events.

Bayesian method improves extreme quantile estimation with zero coverage error.

problem Estimating extreme quantiles with zero coverage error in small samples.
method Bayesian quantile estimation using Jeffreys prior.
result Bayesian method results in zero coverage error, unlike maximum likelihood.

GPDFlow models extreme threshold exceedance with flexible dependence using normalizing flows.

problem Challenges in modeling multivariate threshold exceedance probabilities due to infinite parametrizations.
method GPDFlow uses normalizing flows to flexibly represent dependence without explicit parametric assumptions.
result GPDFlow significantly improves modeling accuracy and flexibility compared to traditional parametric methods.

We investigate the relative information content of six measures of dependence between two random variables XX and YY for large or extreme events for several models of interest for financial time series. The six measures of dependence are respectively the linear correlation ρv+ρ^+_v and Spearman's rho ρs(v)ρ_s(v) conditio…

2002-03-07abs ↗pdf ↗

GenFormer uses deep learning to generate complex stochastic data.

problem Creating synthetic stochastic data that matches real-world statistical properties.
method Transformer-based deep learning model that maps Markov state sequences to time series values.
result GenFormer preserves target marginal distributions and other statistical properties in multivariate spatio-temporal data.

We present a model of credit card profitability, assuming that the card-holder always pays the full outstanding balance. The motivation for the model is to calculate an optimal credit limit, which requires an expression for the expected outstanding balance. We derive its Laplace transform, assuming that purchases are m…

2015-06-17abs ↗pdf ↗

Develops a framework for clustering and distribution matching with bandit feedback.

problem Clustering and distribution matching problems with limited feedback.
method General framework using KK-armed bandit model, Track-and-Stop method, and Frank--Wolfe algorithm.
result Average number of arm pulls matches lower bound, with asymptotic convergence to fundamental limit.

Generative algorithms learn high-dimensional data efficiently and generate new samples.

problem Learning from scarce high-dimensional data.
method Lipschitz-regularized gradient flows and particle-based algorithms.
result Correctly transports gene expression data points with high dimensionality.

Investigates multi-period portfolio optimization for DC plans using buffered Probability of Exceedance.

problem Optimizing long-term Defined Contribution plans with realistic constraints and dynamic dynamics.
method Formulates and solves bilevel optimization problems for pre-commitment and time-consistent Mean-bPoE and Mean-CVaR portfolio optimization.
result Time-consistent Mean-bPoE strategies maintain investor preferences for minimum terminal wealth, unlike Mean-CVaR.

The hidden tail of empirical distributions is analyzed using extreme value theory.

problem Understanding the bias between in-sample mean and true statistical mean for large nn.
method Extreme value theory applied to empirical distributions and their moments.
result The hidden moment of order 0 for power law distributions follows an exponential distribution with expectation 1/n1/n.

Sharp large deviations and Gibbs conditioning for portfolio credit risk models.

problem Analyzing the risk of default in financial portfolios with dependent factors.
method Sharp large deviation estimates and conditional Bahadur-Rao estimates for threshold models with diverging latent factors.
result Conditioned on a large exceedance event, default indicators become asymptotically i.i.d., and loss-given-default is exponentially tilted.

Study on integrability of geodesic flows on Heisenberg group.

problem Integrability of geodesic flows on Heisenberg group.
method Investigation of two classes of normal geodesic flows associated with left-invariant sub-Riemannian metric.
result Left-left configuration is completely integrable in non-commutative sense, while left-right configuration exhibits non-commutative integrability in dimensions > 5.

ES reduces high-probability regret in stochastic linear bandits.

problem High-probability regret in stochastic linear bandits.
method Linear ensemble sampling with standard Gaussian perturbations, analyzing m=Θ(dlogn)m=Θ(d\log n) ensemble size.
result ES achieves ildeO(d3/2n) ilde O(d^{3/2}\sqrt n) high-probability regret, closing the gap to Thompson sampling.

Quantile gradient boosted trees outperform other models in predicting NO2 concentration distributions.

problem Forecasting high NO2 concentration episodes for effective air quality management.
method Compared 10 probabilistic forecasting models for NO2 concentration prediction.
result Quantile gradient boosted trees model outperformed others in predicting NO2 concentration distributions.

Modeling stock returns and volatility using a bivariate gamma generalized Laplace law.

problem Analyzing stock returns and volatility using a new statistical model.
method Maximum likelihood estimation for a bivariate generalized Laplace distribution, simplifying to linear regression.
result Explicit estimators derived with nonstandard convergence rates for certain parameter configurations.

We present KERMIT, a simple insertion-based approach to generative modeling for sequences and sequence pairs. KERMIT models the joint distribution and its decompositions (i.e., marginals and conditionals) using a single neural network and, unlike much prior work, does not rely on a prespecified factorization of the dat…

2019-06-04abs ↗pdf ↗

The statistical properties of the return intervals τqτ_q between successive 1-min volatilities of 30 liquid Chinese stocks exceeding a certain threshold qq are carefully studied. The Kolmogorov-Smirnov (KS) test shows that 12 stocks exhibit scaling behaviors in the distributions of τqτ_q for different thresholds qq. …

2008-07-11abs ↗pdf ↗

Computing the permanent of a non-negative matrix is a core problem with practical applications ranging from target tracking to statistical thermodynamics. However, this problem is also #P-complete, which leaves little hope for finding an exact solution that can be computed efficiently. While the problem admits a fully …

2019-11-26abs ↗pdf ↗

According to the Loss Distribution Approach, the operational risk of a bank is determined as 99.9% quantile of the respective loss distribution, covering unexpected severe events. The 99.9% quantile can be considered a tail event. As supported by the Pickands-Balkema-de Haan Theorem, tail events exceeding some high thr…

2010-12-01abs ↗pdf ↗

Reference class forecasting is a method to remove optimism bias and strategic misrepresentation in infrastructure projects and programmes. In 2012 the Hong Kong government's Development Bureau commissioned a feasibility study on reference class forecasting in Hong Kong - a first for the Asia-Pacific region. This study …

2017-10-03abs ↗pdf ↗

We study the statistical properties of the recurrence intervals ττ between successive trading volumes exceeding a certain threshold qq. The recurrence interval analysis is carried out for the 20 liquid Chinese stocks covering a period from January 2000 to May 2009, and two Chinese indices from January 2003 to April 2…

2010-02-06abs ↗pdf ↗

DCK improves air quality index prediction with probabilistic spatial models.

problem Non-Gaussian, complex spatial structure of air quality index.
method Deep classifier kriging (DCK) for non-Gaussian, nonlinear spatial prediction.
result DCK outperforms conventional methods in predictive accuracy and uncertainty quantification.

We review the recently introduced concept of variety of a financial portfolio and we sketch its importance for risk control purposes. The empirical behaviour of variety, correlation, exceedance correlation and asymmetry of the probability density function of daily returns is discussed. The results obtained are compared…

2001-07-10abs ↗pdf ↗

Paper tackles sampling from non-log-concave distributions using denoising diffusion.

problem Sampling from non-log-concave distributions efficiently.
method DDMC framework, Zeroth-Order Diffusion Monte Carlo (ZOD-MC) algorithm.
result ZOD-MC achieves inverse polynomial dependence on sampling accuracy, efficient for low dimensions.

Financial exchanges provide incentives for limit order book (LOB) liquidity provision to certain market participants, termed designated market makers or designated sponsors. While quoting requirements typically enforce the activity of these participants for a certain portion of the day, we argue that liquidity demand t…

2015-08-18abs ↗pdf ↗

Correctly detecting the semantic type of data columns is crucial for data science tasks such as automated data cleaning, schema matching, and data discovery. Existing data preparation and analysis systems rely on dictionary lookups and regular expression matching to detect semantic types. However, these matching-based …

2019-05-25abs ↗pdf ↗

Let x and y be two (not necessarily distinct) points on a closed Riemannian manifold M of dimension n. According to a celebrated theorem by J.P. Serre there exist infinitely many geodesics between x and y. The length of the shortest of these geodesics is obviously less than the diameter of M. But what can be said about…

2005-12-23abs ↗pdf ↗

The first order behavior of multivariate heavy-tailed random vectors above large radial thresholds is ruled by a limit measure in a regular variation framework. For a high dimensional vector, a reasonable assumption is that the support of this measure is concentrated on a lower dimensional subspace, meaning that certai…

2019-06-26abs ↗pdf ↗

New loss function restores importance weighting in overparameterized models.

problem Restoring importance weighting in overparameterized neural networks.
method Introduced polynomially-tailed losses to restore effects of importance weighting.
result Polynomially-tailed losses improve performance in correcting distribution shift.