Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

72144216288 · Jun 202019922001200920172026
48 results for statistical definition

Differential privacy is a statistical concept that can be explained through hypothesis testing.

problem Formalizing differential privacy as a statistical concept.
method Using David Blackwell's informativeness theorem, the paper shows differential privacy can be understood through hypothesis testing.
result The definition of ff-differential privacy provides a unified framework for analyzing privacy bounds.

The paper explores how market-based returns depend on past trade values.

problem Improving accuracy in forecasting market-based average and volatility of returns.
method Derives the dependence of market-based volatility and higher statistical moments of returns on statistical moments and correlations of current and past trade values.
result Market-based statistical moments can be approximated by a finite number of moments, improving forecast reliability.

Improved Nyström approximation for kernel quadrature with theoretical guarantees.

problem Efficiently approximating positive definite kernels for large datasets.
method Refined sampling and subspace selection in Nyström approximation.
result Novel theoretical guarantees for non-i.i.d. landmark points in kernel quadrature.

Proposes a new fairness definition based on equity for machine learning classification.

problem Machine learning systems can perpetuate societal biases.
method Formalizes a new fairness definition based on equity, operationalizes it for classification, and evaluates its effectiveness.
result Demonstrates the effectiveness of the new fairness definition for equitable classification.

A new imputation method estimates missing values by matching observed marginals from masked data.

problem Missing values in data undermine statistical and machine learning analysis.
method Estimates a distribution from masked observations using positive semi-definite kernel density estimation.
result The method yields both single and multiple imputations from the same fitted density, with statistical consistency and fast adaptive excess risk.

This paper derives radial fields on manifolds of symmetric positive definite matrices.

problem Lack of an expression for radial fields on manifolds of symmetric positive definite matrices.
method Derives an expression for radial fields on manifolds of symmetric positive definite matrices.
result Derives an expression for radial fields on manifolds of symmetric positive definite matrices.

We develope a new and general notion of parametric measure models and statistical models on an arbitrary sample space ΩΩ which does not assume that all measures of the model have the same null sets. This is given by a diffferentiable map from the parameter manifold MM into the set of finite measures or probability me…

2015-10-25abs ↗pdf ↗

A new mechanism for differentially private Fréchet mean on SPD matrices.

problem Privacy-preserving statistical summaries for SPD matrices.
method Tangent Gaussian mechanism for log-Euclidean metric.
result Significantly better utility and computational efficiency.

Differential privacy is a de facto standard in data privacy, with applications in the public and private sectors. A way to explain differential privacy, which is particularly appealing to statistician and social scientists is by means of its statistical hypothesis testing interpretation. Informally, one cannot effectiv…

2019-05-24abs ↗pdf ↗

This survey is an introduction to positive definite kernels and the set of methods they have inspired in the machine learning literature, namely kernel methods. We first discuss some properties of positive definite kernels as well as reproducing kernel Hibert spaces, the natural extension of the set of functions $\{k(x…

2009-11-28abs ↗pdf ↗

Defines a similarity measure for classification distributions.

problem Measuring similarity between classification distributions.
method Proposes task similarity, a novel measure quantifying performance of source distributions on target distributions.
result Empirical task similarity correlates with transfer efficiency and semantic similarity of source distributions.

The paper analyzes Karcher means on restricted PSD matrices with statistical guarantees.

problem Statistical analysis of non-linear manifolds in machine learning.
method Intrinsic mean model on restricted PSD matrices, Karcher mean analysis, extrinsic signal-plus-noise model.
result Non-asymptotic statistical analysis of Karcher means with deterministic error bounds.

We prove in this paper that the weighted volume of the set of integral transportation matrices between two integral histograms r and c of equal sum is a positive definite kernel of r and c when the set of considered weights forms a positive definite matrix. The computation of this quantity, despite being the subject of…

2012-09-12abs ↗pdf ↗

The need for new methods to deal with big data is a common theme in most scientific fields, although its definition tends to vary with the context. Statistical ideas are an essential part of this, and as a partial response, a thematic program on statistical inference, learning, and models in big data was held in 2015 i…

2015-09-09abs ↗pdf ↗

Mathematical foundation for phylogenetic tree uncertainty quantification.

problem Uncertainty in evolutionary relationships between species.
method Introducing the Wald space as a subset of symmetric positive definite matrices, studying its topology and structure, and proposing a new numerical method for geodesics and curvature.
result Wald space has a topology of disjoint open cubes, is contractible, and is a Whitney stratified space of type (A).

Graph-based methods for signal processing have shown promise for the analysis of data exhibiting irregular structure, such as those found in social, transportation, and sensor networks. Yet, though these systems are often dynamic, state-of-the-art methods for signal processing on graphs ignore the dimension of time, tr…

2016-06-22abs ↗pdf ↗

This paper develops copula-based models for forecasting multivariate realized volatility.

problem Forecasting multivariate realized volatility matrices with hidden dependence structure.
method Copula-based time series models to capture hidden dependence structure and ensure positive definiteness.
result Copula-based models achieve significant performance in volatility matrix forecasting.

XAI methods struggle with identifying true predictors from suppressors in linear datasets.

problem XAI methods misidentify suppressor variables as important features.
method Carefully crafted linear ground-truth dataset to study suppressor variables; evaluated various XAI methods.
result Most XAI methods fail to distinguish true predictors from suppressors in linear settings.

Modified relative universality for unbiasedness and consistency in dimension reduction.

problem Gap in proof of unbiasedness and Fisher consistency in relative universality.
method Modified definition of relative universality using ǫ-measurability.
result Established unbiasedness and Fisher consistency rigorously.

AI systems need reliable testing to ensure safety and trustworthiness.

problem Current AI Act lacks functional trustworthiness for AI systems.
method Define technical application distribution, set risk-based performance, and conduct statistically valid testing.
result Reliable functional trustworthiness is essential for AI systems.

Optimal transport (\OT) theory defines a powerful set of tools to compare probability distributions. \OT~suffers however from a few drawbacks, computational and statistical, which have encouraged the proposal of several regularized variants of OT in the recent literature, one of the most notable being the \textit{slice…

2019-02-01abs ↗pdf ↗

Infinite mixture models are commonly used for clustering. One can sample from the posterior of mixture assignments by Monte Carlo methods or find its maximum a posteriori solution by optimization. However, in some problems the posterior is diffuse and it is hard to interpret the sampled partitionings. In this paper, we…

2013-10-01abs ↗pdf ↗

A new network log-ARCH model improves stock market volatility forecasting.

problem Improving stock market volatility forecasting accuracy.
method Dynamic network autoregressive conditional heteroscedasticity (ARCH) model integrating lagged and adjacent node volatility information.
result The model shows significant improvements in forecasting accuracy compared to univariate log-ARCH models.

We examine the out-of-equilibrium phase reported by Plerou {\it et. al.} in Nature, {\bf 421}, 130 (2003) using the data of the New York stock market (NYSE) between the years 2001 --2002. We find that the observed two phase phenomenon is an artifact of the definition of the control parameter coupled with the nature of …

2005-02-15abs ↗pdf ↗

The paper examines how market trade values and volumes affect price autocorrelation.

problem Understanding the impact of market trade values and volumes on price autocorrelation.
method Derives the dependence of price statistical moments and volatility on trade values and volumes, and assesses statistical moments and correlations by conventional frequency-based probabilities.
result Highlights the impact of market trade randomness on price statistical moments and autocorrelation.

Information theory provides principled ways to analyze different inference and learning problems such as hypothesis testing, clustering, dimensionality reduction, classification, among others. However, the use of information theoretic quantities as test statistics, that is, as quantities obtained from empirical data, p…

2012-11-11abs ↗pdf ↗

We introduce a wrapped Gaussian for SPD matrices, enhancing data analysis.

problem Handling circular and non-flat data distributions on SPD manifolds.
method Introduced a non-isotropic wrapped Gaussian using the exponential map, derived theoretical properties, and proposed a maximum likelihood framework.
result Demonstrated the robustness and flexibility of the wrapped Gaussian model on synthetic and real-world datasets.

This paper generalizes beta divergence beyond its classical form associated with power variance functions of Tweedie models. Generalized form is represented by a compact definite integral as a function of variance function of the exponential dispersion model. This compact integral form simplifies derivations of many pr…

2013-06-14abs ↗pdf ↗

Replicable clustering algorithms for k-medians, k-means, and k-centers are proposed.

problem Designing clustering algorithms that produce the same partition on repeated runs under the same distribution.
method Utilizing approximation routines for combinatorial clustering problems in a black-box manner.
result Replicable algorithms for statistical kk-medians, kk-means, and kk-centers with specified approximation and sample complexities.