Simple bounds show most cross-sectional predictability findings are likely true.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study examines local extrema and crossing statistics in financial markets.
Investigates cross-impact kernels for financial asset prices.
The level crossing and inverse statistics analysis of DAX and oil price time series are given. We determine the average frequency of positive-slope crossings, , where is the average waiting time for observing the level again. We estimate the probability , which provides us the probab…
Cross-sectional "Information Coefficient" (IC) is a widely and deeply accepted measure in portfolio management. The paper gives an insight into IC in view of high-dimensional directional statistics: IC is a linear operator on the components of a centralizing-unitizing standardized random vector of next-period cross-sec…
New methods improve cross-conformal prediction's prediction sets without sacrificing coverage guarantees.
Cross-validation is one of the most popular model selection methods in statistics and machine learning. Despite its wide applicability, traditional cross validation methods tend to select overfitting models, due to the ignorance of the uncertainty in the testing sample. We develop a new, statistically principled infere…
Paper introduces statistical learning for point processes.
A well-known issue of Batch Normalization is its significantly reduced effectiveness in the case of small mini-batch sizes. When a mini-batch contains few examples, the statistics upon which the normalization is defined cannot be reliably estimated from it during a training iteration. To address this problem, we presen…
Twinning splits data into fast, statistically similar sets.
In order to emphasize cross-correlations for fluctuations in major market places, series of up and down spins are built from financial data. Patterns frequencies are measured, and statistical tests performed. Strong cross-correlations are emphasized, proving that market moves are collective behaviors.
Analysis of cross-validation for early-stopped gradient descent in high-dimensional regression.
Cross-validation is the workhorse of modern applied statistics and machine learning, as it provides a principled framework for selecting the model that maximizes generalization performance. In this paper, we show that the cross-validation risk is differentiable with respect to the hyperparameters and training data for …
The authors argue against the classification of forecasting methods as machine learning or statistical.
In this paper I show how reliable estimates of the Value of a Statistical Life (VSL) can be obtained using cross sectional data using Garen's instrumental variable (IV) approach. The increase in the range confidence intervals due to the IV setup can be reduced by a factor of 3 by using a proxy to risk attitude. In orde…
We highlight a very simple statistical tool for the analysis of financial bubbles, which has already been studied in [1]. We provide extensive empirical tests of this statistical tool and investigate analytically its link with stocks correlation structure.
A new test statistic speeds up MMD while maintaining power.
DUPLE tackles cross-deployment recognition in fiber-optic perimeter security with meta-learning.
Cross-validation (CV) is a technique for evaluating the ability of statistical models/learning systems based on a given data set. Despite its wide applicability, the rather heavy computational cost can prevent its use as the system size grows. To resolve this difficulty in the case of Bayesian linear regression, we dev…
Study compares cryptocurrency and stock markets using statistical equilibrium models.
Paper uses HPCA for better stock correlation modeling.
Cross-balancing improves causal inference by balancing features with outcome data.
We introduce a new test for detection of power-law cross-correlations among a pair of time series - the rescaled covariance test. The test is based on a power-law divergence of the covariance of the partial sums of the long-range cross-correlated processes. Utilizing a heteroskedasticity and auto-correlation robust est…
With the increasing size of today's data sets, finding the right parameter configuration in model selection via cross-validation can be an extremely time-consuming task. In this paper we propose an improved cross-validation procedure which uses nonparametric testing coupled with sequential analysis to determine the bes…
Novel methods robustify Gromov-Wasserstein distance for cross-domain alignment.
This paper improves model selection with cross-validation using domain knowledge.
MuyGPs efficiently estimates GP hyperparameters using local cross-validation.
Method embeds numeric tabular datasets into a shared vector space for similarity and retrieval.
This paper provides a construction of a quantum statistical mechanical system associated to knots in the 3-sphere and cyclic branched coverings of the 3-sphere, which is an analog, in the sense of arithmetic topology, of the Bost-Connes system, with knots replacing primes, and cyclic branched coverings of the 3-sphere …
Determining the extent to which different cognitive modalities (understood here as the set of cognitive processes underlying the elaboration of a stimulus by the brain) rely on overlapping neural representations is a fundamental issue in cognitive neuroscience. In the last decade, the identification of shared activity …
A new kernel test avoids permutations for independence testing.
We present cross and time series analysis of price fluctuations in the U.S. Treasury fixed income market. By means of techniques borrowed from statistical physics we show that the correlation among bonds depends strongly on the maturity and bonds' price increments do not fulfill the random walk hyphoteses.
In machine learning, statistics, econometrics and statistical physics, cross-validation (CV) is used asa standard approach in quantifying the generalisation performance of a statistical model. A directapplication of CV in time-series leads to the loss of serial correlations, a requirement of preserving anynon-stationar…
Cheap permutation tests speed up distribution testing without sacrificing accuracy.
Statistical arbitrageurs have inelastic demand, contrary to classical models.
As the success of deep learning reaches more grounds, one would like to also envision the potential limits of deep learning. This paper gives a first set of results proving that certain deep learning algorithms fail at learning certain efficiently learnable functions. The results put forward a notion of cross-predictab…
This paper develops dimension-agnostic inference methods for high-dimensional data.
New theory for PCA under weak latent factors, improving inference and testing.
A fast bootstrap method estimates cross-validation standard error.
Statistical machine learning models should be evaluated and validated before putting to work. Conventional k-fold Monte Carlo Cross-Validation (MCCV) procedure uses a pseudo-random sequence to partition instances into k subsets, which usually causes subsampling bias, inflates generalization errors and jeopardizes the r…
TAP transfers knowledge from unlabeled data to improve cross-modal learning.
Better signal detection in undersampled data using joint and cross covariances.
Study finds multifractal cross-correlations between agricultural markets and external uncertainties.
This paper reformulates for better model performance and interpretation.
New tests detect high-order interactions without permutations.
New method controls false discoveries in financial asset pricing.
This study aims to improve communication between fragmented blockchain systems in finance.
While many statistical models and methods are now available for network analysis, resampling network data remains a challenging problem. Cross-validation is a useful general tool for model selection and parameter tuning, but is not directly applicable to networks since splitting network nodes into groups requires delet…