A new non parametric approach to the problem of testing the independence of two random process is developed. The test statistic is the Hilbert Schmidt Independence Criterion (HSIC), which was used previously in testing independence for i.i.d pairs of variables. The asymptotic behaviour of HSIC is established when compu…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We apply multiple testing procedures to the validation of estimated default probabilities in credit rating systems. The goal is to identify rating classes for which the probability of default is estimated inaccurately, while still maintaining a predefined level of committing type I errors as measured by the familywise …
We apply the procedure of Lee et al. to the problem of performing inference on the signal-noise ratio of the asset which displays maximum sample Sharpe ratio over a set of possibly correlated assets. We find a multivariate analogue of the commonly used approximate standard error of the Sharpe ratio to use in this condi…
Proposes a two-stage method for testing variable interactions with FDR control.
Gonogo offers tools for sensitivity experiments in R.
We propose procedures for testing whether stock price processes are martingales based on limit order type betting strategies. We first show that the null hypothesis of martingale property of a stock price process can be tested based on the capital process of a betting strategy. In particular with high frequency Markov …
New test ensures quality of shared data in machine learning.
Proposes a method to estimate and infer networks from multiple high-dimensional point processes.
New methods test discrete distributions faster with local privacy constraints.
Testing independence is of significant interest in many important areas of large-scale inference. Using extreme-value form statistics to test against sparse alternatives and using quadratic form statistics to test against dense alternatives are two important testing procedures for high-dimensional independence. However…
We propose a novel procedure for outlier detection in functional data, in a semi-supervised framework. As the data is functional, we consider the coefficients obtained after projecting the observations onto orthonormal bases (wavelet, PCA). A multiple testing procedure based on the two-sample test is defined in order t…
A density ratio is defined by the ratio of two probability densities. We study the inference problem of density ratios and apply a semi-parametric density-ratio estimator to the two-sample homogeneity test. In the proposed test procedure, the f-divergence between two probability densities is estimated using a density-r…
A method for rank verification in multivariate Gaussian data, improving on existing approaches.
The study identifies extremal dependence in financial markets using a bootstrap-based testing procedure.
A new test assesses how well observed networks fit a specified ERGM model.
Paper introduces a new test for conditional independence using weighted partial copulas.
Post-detection analysis identifies responsible coordinates for multivariate change-points.
Statistical inference based on lossy or incomplete samples is often needed in research areas such as signal/image processing, medical image storage, remote sensing, signal transmission. In this paper, we propose a nonparametric testing procedure based on samples quantized to bits through a computationally efficient…
Wide class of elliptically contoured distributions is a popular model of stock returns distribution. However the important question of adequacy of the model is open. There are some results which reject and approve such model. Such results are obtained by testing some properties of elliptical model for each pair of stoc…
Develops privacy-preserving methods for equivalence testing in healthcare.
We introduce a general non-parametric independence test between right-censored survival times and covariates, which may be multivariate. Our test statistic has a dual interpretation, first in terms of the supremum of a potentially infinite collection of weight-indexed log-rank tests, with weight functions belonging to …
Testing symmetry of a probability distribution is a common question arising from applications in several fields. Particularly, in the study of observables used in the analysis of stock market index variations, the question of symmetry has not been fully investigated by means of statistical procedures. In this work a di…
This study proposes the segmentation procedure of univariate time series based on Fisher's exact test. We show that an adequate change point can be detected as the minimum value of p-value. It is shown that the proposed procedure can detect change points for an artificial time series. We apply the proposed method to fi…
Transforms any test into anytime-valid with sample savings.
New method tests CMI using deep neural networks for high-dimensional data.
New method controls false discoveries in online testing with deadlines.
The paper tests properties of trees in graphical models using covariance queries.
Paper revisits pre-validation method, improving hypothesis testing.
Two tests identify heterogeneous components in distributed learning.
Identifying dependency in multivariate data is a common inference task that arises in numerous applications. However, existing nonparametric independence tests typically require computation that scales at least quadratically with the sample size, making it difficult to apply them to massive data. Moreover, resampling i…
In recent years, many non-traditional classification methods, such as Random Forest, Boosting, and neural network, have been widely used in applications. Their performance is typically measured in terms of classification accuracy. While the classification error rate and the like are important, they do not address a fun…
The paper introduces localized conformal p-values for conditional testing problems.
Multiple hypothesis testing, a situation when we wish to consider many hypotheses, is a core problem in statistical inference that arises in almost every scientific field. In this setting, controlling the false discovery rate (FDR), which is the expected proportion of type I error, is an important challenge for making …
Throughout the last decade, random forests have established themselves as among the most accurate and popular supervised learning methods. While their black-box nature has made their mathematical analysis difficult, recent work has established important statistical properties like consistency and asymptotic normality b…
New test detects independence in streaming data, adapting to data complexity.
We propose the conditional predictive impact (CPI), a consistent and unbiased estimator of the association between one or several features and a given outcome, conditional on a reduced feature set. Building on the knockoff framework of Candès et al. (2018), we develop a novel testing procedure that works in conjunction…
Extends knockoff filter for composite null hypotheses in variable selection.
In recent years, there has been considerable theoretical development regarding variable selection consistency of penalized regression techniques, such as the lasso. However, there has been relatively little work on quantifying the uncertainty in these selection procedures. In this paper, we propose a new method for inf…
We introduce hyppo, a unified library for performing multivariate hypothesis testing, including independence, two-sample, and k-sample testing. While many multivariate independence tests have R packages available, the interfaces are inconsistent and most are not available in Python. hyppo includes many state of the art…
New testing method for robust actor-critic bandit algorithms.
We investigate multiple testing and variable selection using the Least Angle Regression (LARS) algorithm in high dimensions under the assumption of Gaussian noise. LARS is known to produce a piecewise affine solution path with change points referred to as the knots of the LARS path. The key to our results is an express…
Gaussian graphical model is a graphical representation of the dependence structure for a Gaussian random vector. It is recognized as a powerful tool in different applied fields such as bioinformatics, error-control codes, speech language, information retrieval and others. Gaussian graphical model selection is a statist…
Develops abstention procedure for nonparametric regression via variance testing.
Paper introduces detect-then-impute conformal prediction for cellwise outliers.
Study proposes a statistical testing framework for evaluating clustering pipelines.
New test detects differences in heterogeneous datasets.
Test for linearizing 2-input systems with 2D feedback.
Hypothesis testing in the linear regression model is a fundamental statistical problem. We consider linear regression in the high-dimensional regime where the number of parameters exceeds the number of samples (). In order to make informative inference, we assume that the model is approximately sparse, that is th…