Proves Gerber statistic is always non-negative.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Differential privacy is a statistical concept that can be explained through hypothesis testing.
The study defines backdoor detection in ML and proves its infeasibility.
Survey of statistical queries and their applications.
Study CR-statistical submanifolds in holomorphic statistical spaces.
Paper develops PAC verification for hypothesis classes and statistical algorithms.
The paper explores how market-based returns depend on past trade values.
Improved Nyström approximation for kernel quadrature with theoretical guarantees.
Proposes a new fairness definition based on equity for machine learning classification.
We propose a novel framework of the model specification test in regression using unlabeled test data. In many cases, we have conducted statistical inferences based on the assumption that we can correctly specify a model. However, it is difficult to confirm whether a model is correctly specified. To overcome this proble…
We provide a unifying framework linking two classes of statistics used in two-sample and independence testing: on the one hand, the energy distances and distance covariances from the statistics literature; on the other, maximum mean discrepancies (MMD), that is, distances between embeddings of distributions to reproduc…
This work addresses the issue of large covariance matrix estimation in high-dimensional statistical analysis. Recently, improved iterative algorithms with positive-definite guarantee have been developed. However, these algorithms cannot be directly extended to use a nonconvex penalty for sparsity inducing. Generally, a…
A new imputation method estimates missing values by matching observed marginals from masked data.
This paper examines various definitions of adversarial risk and their implications.
This paper derives radial fields on manifolds of symmetric positive definite matrices.
Paper questions RNN and LSTM's long-term memory and introduces a new definition.
We develope a new and general notion of parametric measure models and statistical models on an arbitrary sample space which does not assume that all measures of the model have the same null sets. This is given by a diffferentiable map from the parameter manifold into the set of finite measures or probability me…
A new mechanism for differentially private Fréchet mean on SPD matrices.
Differential privacy is a de facto standard in data privacy, with applications in the public and private sectors. A way to explain differential privacy, which is particularly appealing to statistician and social scientists is by means of its statistical hypothesis testing interpretation. Informally, one cannot effectiv…
Improved method for computing Fréchet means on SPD matrices.
This survey is an introduction to positive definite kernels and the set of methods they have inspired in the machine learning literature, namely kernel methods. We first discuss some properties of positive definite kernels as well as reproducing kernel Hibert spaces, the natural extension of the set of functions $\{k(x…
Defines a similarity measure for classification distributions.
The paper uses statistics to improve the explainability of models.
The paper analyzes Karcher means on restricted PSD matrices with statistical guarantees.
We prove in this paper that the weighted volume of the set of integral transportation matrices between two integral histograms r and c of equal sum is a positive definite kernel of r and c when the set of considered weights forms a positive definite matrix. The computation of this quantity, despite being the subject of…
Extends differential privacy to Riemannian manifolds, improving utility.
The need for new methods to deal with big data is a common theme in most scientific fields, although its definition tends to vary with the context. Statistical ideas are an essential part of this, and as a partial response, a thematic program on statistical inference, learning, and models in big data was held in 2015 i…
Mathematical foundation for phylogenetic tree uncertainty quantification.
Graph-based methods for signal processing have shown promise for the analysis of data exhibiting irregular structure, such as those found in social, transportation, and sensor networks. Yet, though these systems are often dynamic, state-of-the-art methods for signal processing on graphs ignore the dimension of time, tr…
This paper develops copula-based models for forecasting multivariate realized volatility.
In machine learning or statistics, it is often desirable to reduce the dimensionality of a sample of data points in a high dimensional space . This paper introduces a dimensionality reduction method where the embedding coordinates are the eigenvectors of a positive semi-definite kernel obtained as the sol…
XAI methods struggle with identifying true predictors from suppressors in linear datasets.
Modified relative universality for unbiasedness and consistency in dimension reduction.
We describe society as a nonequilibrium probabilistic system: N individuals occupy W resource states in it and produce entropy S over definite time periods. Resulting thermodynamics is however unusual because a second entropy, H, measures a typically social feature, inequality or diversity in the distribution of availa…
AI systems need reliable testing to ensure safety and trustworthiness.
Optimal transport (\OT) theory defines a powerful set of tools to compare probability distributions. \OT~suffers however from a few drawbacks, computational and statistical, which have encouraged the proposal of several regularized variants of OT in the recent literature, one of the most notable being the \textit{slice…
Infinite mixture models are commonly used for clustering. One can sample from the posterior of mixture assignments by Monte Carlo methods or find its maximum a posteriori solution by optimization. However, in some problems the posterior is diffuse and it is hard to interpret the sampled partitionings. In this paper, we…
Financial networks analyzed using statistical physics methods.
Bagging is a device intended for reducing the prediction error of learning algorithms. In its simplest form, bagging draws bootstrap samples from the training sample, applies the learning algorithm to each bootstrap sample, and then averages the resulting prediction rules. We extend the definition of bagging from stati…
A new network log-ARCH model improves stock market volatility forecasting.
We examine the out-of-equilibrium phase reported by Plerou {\it et. al.} in Nature, {\bf 421}, 130 (2003) using the data of the New York stock market (NYSE) between the years 2001 --2002. We find that the observed two phase phenomenon is an artifact of the definition of the control parameter coupled with the nature of …
The paper examines how market trade values and volumes affect price autocorrelation.
This paper proposes a new integrated variance estimator based on order statistics within the framework of jump-diffusion models. Its ability to disentangle the integrated variance from the total process quadratic variation is confirmed by both simulated and empirical tests. For practical purposes, we introduce an itera…
Information theory provides principled ways to analyze different inference and learning problems such as hypothesis testing, clustering, dimensionality reduction, classification, among others. However, the use of information theoretic quantities as test statistics, that is, as quantities obtained from empirical data, p…
We introduce a wrapped Gaussian for SPD matrices, enhancing data analysis.
This paper generalizes beta divergence beyond its classical form associated with power variance functions of Tweedie models. Generalized form is represented by a compact definite integral as a function of variance function of the exponential dispersion model. This compact integral form simplifies derivations of many pr…
Replicable clustering algorithms for k-medians, k-means, and k-centers are proposed.
New distances measure mixtures of Gaussians, useful in machine learning.