Enhanced KS test detects tail differences more sensitively.
problem Detecting differences in data tail regions.
method Higher-order IPM defined over a total variation ball.
result Asymptotic null distribution for the test.
Accurate goodness-of-fit tests for the extreme tails of empirical distributions is a very important issue, relevant in many contexts, including geophysics, insurance, and finance. We have derived exact asymptotic results for a generalization of the large-sample Kolmogorov-Smirnov test, well suited to testing these extr…
New test uses neural networks to compare distributions, outperforming traditional methods.
problem Comparing distributions in high dimensions and higher orders of smoothness.
method Integral probability metrics with Radon bounded variation functions and neural networks.
result The Radon-Kolmogorov-Smirnov (RKS) test outperforms traditional methods in distinguishing distributions.
Test to distinguish metric space distributions using pairwise distances.
problem Distinguishing between two distributions on a metric space.
method Use pairwise distances and a two-sample Kolmogorov--Smirnov test.
result Can determine if two distributions are the same using finite data.
KSGAN uses KS distance for deep generative modeling.
problem Deep generative modeling challenges, especially for multivariate distributions.
method Formulates adversarial training as minimization of KS distance, using quantile function as critic.
result KSGAN trained distributions closely match target distributions.
Proposes a new TS algorithm for non-stationary bandits using KS tests.
problem Non-stationary multi-armed bandit problems.
method Active detection of change points using KS tests and adaptive Thompson Sampling.
result Sub-linear regret demonstrated for the two-armed bandit case.
Study evaluates two-sample tests for validating generative models in high dimensions.
problem Validating the performance and efficiency of non-parametric two-sample tests for high-dimensional generative models.
method Proposes and evaluates the sliced Wasserstein distance, mean of Kolmogorov-Smirnov statistics, and novel sliced Kolmogorov-Smirnov statistic.
result One-dimensional-based tests provide comparable sensitivity to other multivariate metrics but with lower computational cost.
SurvLIME-KS improves survival model explanations robustly.
problem Improving explanations of unreliable survival models.
method SurvLIME-KS combines Cox proportional hazards model and Kolmogorov-Smirnov bounds for robust optimization.
result SurvLIME-KS minimizes average distance and maximizes distance in approximating cumulative hazard functions.
The paper introduces a spline-based method for calibrating neural networks.
problem Ensuring neural network outputs are reliable for safety-critical applications.
method Approximating the empirical cumulative distribution function using splines to map network outputs to calibrated probabilities.
result The spline-based recalibration consistently outperforms existing methods on calibration measures.
This paper tackles G-ZSL by learning compositional spaces to classify unseen classes.
problem Classifying unseen classes in a test set.
method Space decomposition method to estimate and fine-tune decision boundaries between source and target classes.
result State-of-the-art performance on multiple G-ZSL benchmarks.
MAGDiff detects data shifts in neural networks without retraining.
problem Neural networks' sensitivity to data distribution shifts.
method Extracts MAGDiff representations from neural networks to detect shifts.
result MAGDiff representations improve data set shift detection.
Many problems in finance are related to first passage times. Among all of them, we chose three on which we contributed personally. Our first example relates Kolmogorov-Smirnov like goodness-of-fit tests, modified in such a way that tail events and core events contribute equally to the test (in the standard Kolmogorov-S…
The paper introduces a new method to detect rough volatility and market states using fractional derivatives.
problem Testing self-similarity in fractional processes from a single observed trajectory is difficult under long-range dependence.
method The paper introduces a regime-adaptive KS/GL--KS framework based on the discrete Grünwald--Letnikov (GL) fractional derivative.
result The method detects rough volatility and persistent, anti-persistent, or efficient market states in financial applications.
The paper fits a seven-parameter GTS distribution to financial data.
problem Nonexistence of GTS probability density function makes MLE inadequate.
method Used fractional Fourier transform to circumvent MLE and provide good parameter estimation.
result The GTS distribution fits financial data significantly better than other models.
We present a novel modulation level classification (MLC) method based on probability distribution distance functions. The proposed method uses modified Kuiper and Kolmogorov-Smirnov distances to achieve low computational complexity and outperforms the state of the art methods based on cumulants and goodness-of-fit test…
We revisit the Kolmogorov-Smirnov and Cramér-von Mises goodness-of-fit (GoF) tests and propose a generalisation to identically distributed, but dependent univariate random variables. We show that the dependence leads to a reduction of the "effective" number of independent observations. The generalised GoF tests are not…
KS(conf) detects when ConvNets operate outside their training specifications.
problem Detecting when ConvNets operate outside their training data distribution.
method Applying a Kolmogorov-Smirnov test to predicted confidence values.
result KS(conf) reliably detects out-of-specs situations with minimal overhead.
We study a novel spline-like basis, which we name the "falling factorial basis", bearing many similarities to the classic truncated power basis. The advantage of the falling factorial basis is that it enables rapid, linear-time computations in basis matrix multiplication and basis matrix inversion. The falling factoria…
We investigate the probability distribution of the return intervals τ between successive 1-min volatilities of two Chinese indices exceeding a certain threshold q. The Kolmogorov-Smirnov (KS) tests show that the two indices exhibit multiscaling behavior in the distribution of τ, which follows a stretched exponent…
Nonparametric two sample or homogeneity testing is a decision theoretic problem that involves identifying differences between two random variables without making parametric assumptions about their underlying distributions. The literature is old and rich, with a wide variety of statistics having being intelligently desi…
Estimates Hurst exponent of log-volatility using KS statistic, addressing serial correlation in financial data.
problem Estimating Hurst exponent of log-volatility in financial time series with serial correlation.
method Proposes a random permutation procedure to remove serial correlation, using the Kolmogorov-Smirnov statistic for distribution-based estimation.
result Establishes the asymptotic variance of the estimator and reveals statistically significant hierarchy of roughness in volatility measures.
The statistical properties of the return intervals τq between successive 1-min volatilities of 30 liquid Chinese stocks exceeding a certain threshold q are carefully studied. The Kolmogorov-Smirnov (KS) test shows that 12 stocks exhibit scaling behaviors in the distributions of τq for different thresholds q. …
We compute the analytic expression of the probability distributions F{AEX,+} and F{AEX,-} of the normalized positive and negative AEX (Netherlands) index daily returns r(t). Furthermore, we define the αre-scaled AEX daily index positive returns r(t)^αand negative returns (-r(t))^αthat we call, after normalization, the …
This paper reports empirical evidence that a neural networks model is applicable to the statistically reliable prediction of foreign exchange rates. Time series data and technical indicators such as moving average, are fed to neural nets to capture the underlying "rules" of the movement in currency exchange rates. The …
Hypothesis tests in models whose dimension far exceeds the sample size can be formulated much like the classical studentized tests only after the initial bias of estimation is removed successfully. The theory of debiased estimators can be developed in the context of quantile regression models for a fixed quantile value…
We compute the analytic expression of the probability distributions F{FTSE100,+} and F{FTSE100,-} of the normalized positive and negative FTSE100 (UK) index daily returns r(t). Furthermore, we define the alpha re-scaled FTSE100 daily index positive returns r(t)^alpha and negative returns (-r(t))^alpha that we call, aft…
Unified score and distance-based GoF tests for model adequacy.
problem Difficulty in extending score-based GoF tests to nonparametric alternatives.
method Introducing semiparametric kernelized Stein discrepancy (SKSD) test.
result SKSD test is computationally efficient and universally consistent.
Improved change point detection using matched filters for non-parametric tests.
problem False positives and localization ambiguity in non-parametric two-sample tests.
method Derived and applied matched filters for various two-sample tests.
result Matched filters reduce false positives and improve test precision.
In terms of the stock exchange returns, we compute the analytic expression of the probability distributions F{DAX,+} and F{DAX,-} of the normalized positive and negative DAX (Germany) index daily returns r(t). Furthermore, we define the alpha re-scaled DAX daily index positive returns r(t)^alpha and negative returns (-…
A new UU-test decides unimodality of datasets.
problem Deciding on the unimodality of a dataset for better data analysis.
method UU-test operates on the empirical cumulative density function (ecdf) to build a piecewise linear approximation that models the data as a Uniform Mixture Model.
result The UU-test provides a statistical model of the data in the form of a Uniform Mixture Model.
This article proposes a method to quantify the structure of a bipartite graph using a network entropy per link. The network entropy of a bipartite graph with random links is calculated both numerically and theoretically. As an application of the proposed method to analyze collective behavior, the affairs in which parti…
We investigate the probability distributions of the recurrence intervals τ between consecutive 1-min returns above a positive threshold q>0 or below a negative threshold q<0 of two indices and 20 individual stocks in China's stock market. The distributions of recurrence intervals for positive and negative thresho…
A new method, InfoGuide, improves automatic clustering analysis.
problem Lack of automatic clustering analysis frameworks.
method Capturing traces of information gain between clustering retrievals.
result InfoGuide can enable more automatic clustering analysis.
agtboost speeds up gradient tree boosting with automatic complexity adjustment.
problem Speeding up and simplifying gradient tree boosting computations.
method Adaptive gradient tree boosting with automatic complexity adjustment and feature importance.
result Significant decrease in computation time and simplification of model complexity.
We perform return interval analysis of 1-min {\em{realized volatility}} defined by the sum of absolute high-frequency intraday returns for the Shanghai Stock Exchange Composite Index (SSEC) and 22 constituent stocks of SSEC. The scaling behavior and memory effect of the return intervals between successive realized vola…
This study compares different types of normalizing flows for generating complex distributions.
problem Comparing different types of normalizing flows for generating complex distributions.
method Real-valued non-Volume preserving (RealNVP), masked autoregressive flow (MAF), coupling rational quadratic spline (C-RQS), and autoregressive rational quadratic spline (A-RQS) were compared using statistical tests.
result A-RQS algorithm outperforms others in terms of accuracy and training speed.
The paper proposes a framework to calibrate multi-agent simulation models from output series using Bayesian optimization.
problem Calibrating multi-agent simulation models from observable output series.
method Novel eligibility set concept, two-sample Kolmogorov-Smirnov test with Bonferroni correction, Bayesian optimization (BO), and trust-region BO (TuRBO).
result Demonstrated the efficiency of the proposed framework using numerical experiments.
Two methods improve Gaussian process predictive distributions' calibration.
problem Improving the reliability of Gaussian process predictive intervals.
method Introduces two methods: cps-gp and bcr-gp, both adapting conformal predictive systems to GP interpolation.
result Both methods provide finite-sample marginal calibration and smooth predictive distributions.
Estimates financial market impacts of COVID-19 using time-varying kernel density.
problem Estimating the impact of COVID-19 on financial markets over time.
method Time-varying kernel density estimation with Kolmogorov-Smirnov statistic.
result Determines the chronology and regional disparities of financial market impacts.
Generative Adversarial Networks improve robust statistics for various distributions.
problem Estimating unknown parameters in adversarially corrupted samples.
method Designing GANs with specific loss functions for robust estimation.
result Extends robust estimation to broader families of distributions.
Model predicts daily closing price distributions in call auctions.
problem Predicting price distributions in financial markets.
method Modeling price formation in call auctions with random orders and equilibrium equation.
result Model accurately predicts daily closing price distributions for financial indices.
We study the statistical properties of the recurrence intervals τ between successive trading volumes exceeding a certain threshold q. The recurrence interval analysis is carried out for the 20 liquid Chinese stocks covering a period from January 2000 to May 2009, and two Chinese indices from January 2003 to April 2…
Study learns optimal auctions from corrupted or perturbed bidder valuation samples.
problem Learning revenue-optimal auctions from corrupted or perturbed samples.
method Proves upper bounds, proposes algorithms for learning near-optimal auctions.
result Proves tight upper bounds and proposes algorithms for near-optimal auctions.
GBST model improves credit risk quantification using survival analysis.
problem Quantifying credit risk in heterogeneous consumer finance data.
method Gradient boosting survival tree (GBST) model integrating survival analysis and gradient boosting.
result GBST model outperforms existing survival models in credit risk quantification.
New tensor-based method for estimating stock correlation matrices.
problem Choosing a proper sample period for estimating correlation matrices.
method Slice-Diagonal Tensor (SDT) factorization technique.
result The new method produces a stable correlation matrix unaffected by the sample period.
Markov chain decoders improve generative models' ability to produce heavy-tailed data.
problem Generative models struggle with heavy-tailed distributions.
method Replaced Gaussian decoder with Markov chain-based Phase-Type distributions.
result Significantly reduced tail Kolmogorov-Smirnov distance and extreme quantile error.
RG-TTA adapts neural forecasters to streaming time series shifts by modulating adaptation intensity.
problem Adapting neural forecasters to distribution shifts in streaming time series data.
method RG-TTA uses a meta-controller that continuously modulates adaptation intensity based on distributional similarity.
result RG-TTA achieves the lowest MSE in 156 of 224 seed-averaged experiments, reducing MSE by 5.7% vs TTA.
Novel anti-grokking phase discovered in neural networks, revealed by HTSR layer quality metric.
problem Understanding and detecting overfitting in neural networks.
method 3-layer MLP, weight decay, HTSR layer quality metric α, correlation traps, Kolmogorov–Smirnov tests. result Anti-grokking phase detected late in training, revealed by α<2 and correlation traps.