Accurate goodness-of-fit tests for the extreme tails of empirical distributions is a very important issue, relevant in many contexts, including geophysics, insurance, and finance. We have derived exact asymptotic results for a generalization of the large-sample Kolmogorov-Smirnov test, well suited to testing these extr…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New test uses neural networks to compare distributions, outperforming traditional methods.
We present an extension of the Kolmogorov-Smirnov (KS) two-sample test, which can be more sensitive to differences in the tails. Our test statistic is an integral probability metric (IPM) defined over a higher-order total variation ball, recovering the original KS test as its simplest case. We give an exact representer…
KSGAN uses KS distance for deep generative modeling.
Proposes a new TS algorithm for non-stationary bandits using KS tests.
Study evaluates two-sample tests for validating generative models in high dimensions.
SurvLIME-KS improves survival model explanations robustly.
The paper introduces a spline-based method for calibrating neural networks.
MAGDiff detects data shifts in neural networks without retraining.
We study the problem of distinguishing between two distributions on a metric space; i.e., given metric measure spaces and , we are interested in the problem of determining from finite data whether or not is . The key is to use pairwise distances between observat…
Many problems in finance are related to first passage times. Among all of them, we chose three on which we contributed personally. Our first example relates Kolmogorov-Smirnov like goodness-of-fit tests, modified in such a way that tail events and core events contribute equally to the test (in the standard Kolmogorov-S…
The paper introduces a new method to detect rough volatility and market states using fractional derivatives.
The paper fits a seven-parameter GTS distribution to financial data.
We present a novel modulation level classification (MLC) method based on probability distribution distance functions. The proposed method uses modified Kuiper and Kolmogorov-Smirnov distances to achieve low computational complexity and outperforms the state of the art methods based on cumulants and goodness-of-fit test…
We revisit the Kolmogorov-Smirnov and Cramér-von Mises goodness-of-fit (GoF) tests and propose a generalisation to identically distributed, but dependent univariate random variables. We show that the dependence leads to a reduction of the "effective" number of independent observations. The generalised GoF tests are not…
We study a novel spline-like basis, which we name the "falling factorial basis", bearing many similarities to the classic truncated power basis. The advantage of the falling factorial basis is that it enables rapid, linear-time computations in basis matrix multiplication and basis matrix inversion. The falling factoria…
We investigate the probability distribution of the return intervals between successive 1-min volatilities of two Chinese indices exceeding a certain threshold . The Kolmogorov-Smirnov (KS) tests show that the two indices exhibit multiscaling behavior in the distribution of , which follows a stretched exponent…
Nonparametric two sample or homogeneity testing is a decision theoretic problem that involves identifying differences between two random variables without making parametric assumptions about their underlying distributions. The literature is old and rich, with a wide variety of statistics having being intelligently desi…
Estimates Hurst exponent of log-volatility using KS statistic, addressing serial correlation in financial data.
The statistical properties of the return intervals between successive 1-min volatilities of 30 liquid Chinese stocks exceeding a certain threshold are carefully studied. The Kolmogorov-Smirnov (KS) test shows that 12 stocks exhibit scaling behaviors in the distributions of for different thresholds . …
We compute the analytic expression of the probability distributions F{AEX,+} and F{AEX,-} of the normalized positive and negative AEX (Netherlands) index daily returns r(t). Furthermore, we define the αre-scaled AEX daily index positive returns r(t)^αand negative returns (-r(t))^αthat we call, after normalization, the …
This paper reports empirical evidence that a neural networks model is applicable to the statistically reliable prediction of foreign exchange rates. Time series data and technical indicators such as moving average, are fed to neural nets to capture the underlying "rules" of the movement in currency exchange rates. The …
Hypothesis tests in models whose dimension far exceeds the sample size can be formulated much like the classical studentized tests only after the initial bias of estimation is removed successfully. The theory of debiased estimators can be developed in the context of quantile regression models for a fixed quantile value…
We compute the analytic expression of the probability distributions F{FTSE100,+} and F{FTSE100,-} of the normalized positive and negative FTSE100 (UK) index daily returns r(t). Furthermore, we define the alpha re-scaled FTSE100 daily index positive returns r(t)^alpha and negative returns (-r(t))^alpha that we call, aft…
Unified score and distance-based GoF tests for model adequacy.
This paper studies the problem of Generalized Zero-shot Learning (G-ZSL), whose goal is to classify instances belonging to both seen and unseen classes at the test time. We propose a novel space decomposition method to solve G-ZSL. Some previous models with space decomposition operations only calibrate the confident pr…
Improved change point detection using matched filters for non-parametric tests.
A new UU-test decides unimodality of datasets.
In terms of the stock exchange returns, we compute the analytic expression of the probability distributions F{DAX,+} and F{DAX,-} of the normalized positive and negative DAX (Germany) index daily returns r(t). Furthermore, we define the alpha re-scaled DAX daily index positive returns r(t)^alpha and negative returns (-…
This article proposes a method to quantify the structure of a bipartite graph using a network entropy per link. The network entropy of a bipartite graph with random links is calculated both numerically and theoretically. As an application of the proposed method to analyze collective behavior, the affairs in which parti…
We investigate the probability distributions of the recurrence intervals between consecutive 1-min returns above a positive threshold or below a negative threshold of two indices and 20 individual stocks in China's stock market. The distributions of recurrence intervals for positive and negative thresho…
A new method, InfoGuide, improves automatic clustering analysis.
agtboost speeds up gradient tree boosting with automatic complexity adjustment.
We perform return interval analysis of 1-min {\em{realized volatility}} defined by the sum of absolute high-frequency intraday returns for the Shanghai Stock Exchange Composite Index (SSEC) and 22 constituent stocks of SSEC. The scaling behavior and memory effect of the return intervals between successive realized vola…
Credit scoring plays a vital role in the field of consumer finance. Survival analysis provides an advanced solution to the credit-scoring problem by quantifying the probability of survival time. In order to deal with highly heterogeneous industrial data collected in Chinese market of consumer finance, we propose a nonp…
This study compares different types of normalizing flows for generating complex distributions.
Two methods improve Gaussian process predictive distributions' calibration.
The paper proposes a framework to calibrate multi-agent simulation models from output series using Bayesian optimization.
Estimates financial market impacts of COVID-19 using time-varying kernel density.
Generative Adversarial Networks improve robust statistics for various distributions.
Computer vision systems for automatic image categorization have become accurate and reliable enough that they can run continuously for days or even years as components of real-world commercial applications. A major open problem in this context, however, is quality control. Good classification performance can only be ex…
We study the statistical properties of the recurrence intervals between successive trading volumes exceeding a certain threshold . The recurrence interval analysis is carried out for the 20 liquid Chinese stocks covering a period from January 2000 to May 2009, and two Chinese indices from January 2003 to April 2…
Study learns optimal auctions from corrupted or perturbed bidder valuation samples.
New tensor-based method for estimating stock correlation matrices.
Markov chain decoders improve generative models' ability to produce heavy-tailed data.
RG-TTA adapts neural forecasters to streaming time series shifts by modulating adaptation intensity.
Novel anti-grokking phase discovered in neural networks, revealed by HTSR layer quality metric.
Principal component analysis (PCA) is very popular to perform dimension reduction. The selection of the number of significant components is essential but often based on some practical heuristics depending on the application. Only few works have proposed a probabilistic approach able to infer the number of significant c…