Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

82164245327 · May 202619922001200920172026
48 results for crossing statistics

Simple bounds show most cross-sectional predictability findings are likely true.

problem Determining the validity of cross-sectional return predictability findings.
method Developed simple and intuitive bounds on the false discovery rate (FDR).
result Bounds show the FDR is small, indicating most findings are likely true.

Study examines local extrema and crossing statistics in financial markets.

problem Understanding local extrema and crossing statistics in financial markets.
method Excursion set theory, numerical computation, theoretical prediction, clustering of geometrical measures, cross-correlation, Singular Value Decomposition.
result Excursion sets reveal statistical coherency and sensitivity to crises in financial markets.

Investigates cross-impact kernels for financial asset prices.

problem Understanding and parameterizing cross-impact kernels for financial asset prices.
method Examined martingale-admissible and no-statistical-arbitrage-admissible kernels, determined their overlap, and provided calibration formulas.
result Identified the overlap between martingale-admissible and no-statistical-arbitrage-admissible kernels and provided formulas for their calibration.

The level crossing and inverse statistics analysis of DAX and oil price time series are given. We determine the average frequency of positive-slope crossings, να+ν_α^+, where Tα=1/να+T_α =1/ν_α^+ is the average waiting time for observing the level αα again. We estimate the probability P(K,α)P(K, α), which provides us the probab…

2010-01-25abs ↗pdf ↗

New methods improve cross-conformal prediction's prediction sets without sacrificing coverage guarantees.

problem Improving the width of prediction sets in cross-conformal prediction.
method Proposed new variants of existing methods based on recent results on more efficient combination of p-values.
result Smaller prediction sets achieved without compromising theoretical guarantees.

Cross-validation is one of the most popular model selection methods in statistics and machine learning. Despite its wide applicability, traditional cross validation methods tend to select overfitting models, due to the ignorance of the uncertainty in the testing sample. We develop a new, statistically principled infere…

2017-03-23abs ↗pdf ↗

A well-known issue of Batch Normalization is its significantly reduced effectiveness in the case of small mini-batch sizes. When a mini-batch contains few examples, the statistics upon which the normalization is defined cannot be reliably estimated from it during a training iteration. To address this problem, we presen…

2020-02-13abs ↗pdf ↗

In order to emphasize cross-correlations for fluctuations in major market places, series of up and down spins are built from financial data. Patterns frequencies are measured, and statistical tests performed. Strong cross-correlations are emphasized, proving that market moves are collective behaviors.

2000-01-20abs ↗pdf ↗

Analysis of cross-validation for early-stopped gradient descent in high-dimensional regression.

problem Inconsistency of GCV for early-stopped GD in high-dimensional least squares regression.
method Theoretical analysis of GCV and LOOCV applied to early-stopped GD in high-dimensional least squares regression.
result LOOCV converges uniformly to the prediction risk of early-stopped GD, while GCV is generically inconsistent.

The authors argue against the classification of forecasting methods as machine learning or statistical.

problem The classification of forecasting methods as machine learning or statistical limits insights into their appropriateness and effectiveness.
method Alternative characteristics of forecasting methods are proposed to draw meaningful conclusions.
result The distinction between machine learning and statistical forecasting methods is not fundamental.

DUPLE tackles cross-deployment recognition in fiber-optic perimeter security with meta-learning.

problem Cross-deployment recognition challenges in fiber-optic perimeter security due to label scarcity and distribution shifts.
method DUPLE employs statistically guided meta-learning to enhance recognition robustness across unseen deployments.
result DUPLE consistently outperforms traditional and meta-learning baselines in cross-deployment DFOS benchmarks.

Cross-validation (CV) is a technique for evaluating the ability of statistical models/learning systems based on a given data set. Despite its wide applicability, the rather heavy computational cost can prevent its use as the system size grows. To resolve this difficulty in the case of Bayesian linear regression, we dev…

2016-10-25abs ↗pdf ↗

Study compares cryptocurrency and stock markets using statistical equilibrium models.

problem Comparing the stochastic structure of cryptocurrency and stock markets.
method Applied QRSE model to analyze daily returns of cryptocurrencies and S&P 500 companies.
result Revealed differences in informational efficiency between cryptocurrency and stock markets.

Cross-balancing improves causal inference by balancing features with outcome data.

problem Balancing features for valid causal inference when outcome data is available.
method Cross-balancing using sample splitting to separate feature construction and weight estimation errors.
result Cross-balancing produces consistent, asymptotically normal, and efficient estimators under mild conditions.

We introduce a new test for detection of power-law cross-correlations among a pair of time series - the rescaled covariance test. The test is based on a power-law divergence of the covariance of the partial sums of the long-range cross-correlated processes. Utilizing a heteroskedasticity and auto-correlation robust est…

2013-07-17abs ↗pdf ↗

With the increasing size of today's data sets, finding the right parameter configuration in model selection via cross-validation can be an extremely time-consuming task. In this paper we propose an improved cross-validation procedure which uses nonparametric testing coupled with sequential analysis to determine the bes…

2012-06-11abs ↗pdf ↗

This paper improves model selection with cross-validation using domain knowledge.

problem Improving model selection with cross-validation risk estimation.
method Establishes distribution-free deviation bounds using VC dimension, formalizes Learning Spaces based on domain knowledge.
result Enhanced generalization through selection of candidate models based on domain knowledge.

MuyGPs efficiently estimates GP hyperparameters using local cross-validation.

problem Efficiently estimating GP hyperparameters for large datasets.
method Uses nearest neighbors structure and leave-one-out cross-validation.
result Outperforms state-of-the-art competitors in time and prediction accuracy.

Method embeds numeric tabular datasets into a shared vector space for similarity and retrieval.

problem Lack of meaningful representation for numeric tabular datasets in large language models.
method Structured exploratory data analysis descriptors, sentence transformer embedding, CCA for cross-dataset alignment.
result Total P@1 score of 0.9 across 15 datasets, robust nearest-neighbor retrieval and cluster structure.

This paper provides a construction of a quantum statistical mechanical system associated to knots in the 3-sphere and cyclic branched coverings of the 3-sphere, which is an analog, in the sense of arithmetic topology, of the Bost-Connes system, with knots replacing primes, and cyclic branched coverings of the 3-sphere …

2016-02-16abs ↗pdf ↗

Determining the extent to which different cognitive modalities (understood here as the set of cognitive processes underlying the elaboration of a stimulus by the brain) rely on overlapping neural representations is a fundamental issue in cognitive neuroscience. In the last decade, the identification of shared activity …

2019-10-08abs ↗pdf ↗

We present cross and time series analysis of price fluctuations in the U.S. Treasury fixed income market. By means of techniques borrowed from statistical physics we show that the correlation among bonds depends strongly on the maturity and bonds' price increments do not fulfill the random walk hyphoteses.

2000-03-02abs ↗pdf ↗

In machine learning, statistics, econometrics and statistical physics, cross-validation (CV) is used asa standard approach in quantifying the generalisation performance of a statistical model. A directapplication of CV in time-series leads to the loss of serial correlations, a requirement of preserving anynon-stationar…

2019-10-21abs ↗pdf ↗

As the success of deep learning reaches more grounds, one would like to also envision the potential limits of deep learning. This paper gives a first set of results proving that certain deep learning algorithms fail at learning certain efficiently learnable functions. The results put forward a notion of cross-predictab…

2018-12-16abs ↗pdf ↗

This paper develops dimension-agnostic inference methods for high-dimensional data.

problem Understanding how classical inference methods behave in high-dimensional settings.
method Using variational representations, sample splitting, and self-normalization to create a refined test statistic.
result The resulting statistic has a Gaussian limiting distribution regardless of how dimensionality scales with sample size.

New theory for PCA under weak latent factors, improving inference and testing.

problem Statistical inference for PCA with weak latent factors and cross-sectional dependence.
method Comprehensive estimation and inference theory for PCA under nearly minimal factor strength, non-asymptotic.
result Asymptotic normality of PCA-based estimator for NTN\asymp T with SNR growth rate.

Statistical machine learning models should be evaluated and validated before putting to work. Conventional k-fold Monte Carlo Cross-Validation (MCCV) procedure uses a pseudo-random sequence to partition instances into k subsets, which usually causes subsampling bias, inflates generalization errors and jeopardizes the r…

2019-07-04abs ↗pdf ↗

TAP transfers knowledge from unlabeled data to improve cross-modal learning.

problem Improving supervised learning performance using unlabeled data from a different modality.
method Probabilistic approach for missing information estimation, kernel regression, cross-attention module, TAP neural network.
result TAP significantly improves generalization across different domains and neural network architectures.

Better signal detection in undersampled data using joint and cross covariances.

problem Detecting shared signals in high-dimensional data with limited samples.
method Analysis of three covariance matrices: individual, cross, and joint.
result Joint and cross covariance matrices detect signals earlier than individual covariances.

Study finds multifractal cross-correlations between agricultural markets and external uncertainties.

problem Investigating relationships between agricultural spot markets and external uncertainties.
method Multifractal detrending moving-average cross-correlation analysis (MF-X-DMA).
result Maize exhibits intrinsic joint multifractality with all uncertainty proxies.

This paper reformulates FβF_β for better model performance and interpretation.

problem Optimizing model performance and interpretation using FβF_β metric.
method Reformulate FβF_β metric to facilitate statistical distributions and dynamic penalty weights.
result Better and interpretable results with a 14% boost in F1F_1 score for IMDB data.

This study aims to improve communication between fragmented blockchain systems in finance.

problem Inefficient and insecure communication in fragmented blockchain systems.
method Analysis of cross-chain interoperability protocols and their properties.
result Comparison and evaluation of cross-chain interoperability protocols.

While many statistical models and methods are now available for network analysis, resampling network data remains a challenging problem. Cross-validation is a useful general tool for model selection and parameter tuning, but is not directly applicable to networks since splitting network nodes into groups requires delet…

2016-12-14abs ↗pdf ↗