Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

101202303404 · Jun 202019922001200920172026
48 results for statistical dimension

This paper develops dimension-agnostic inference methods for high-dimensional data.

problem Understanding how classical inference methods behave in high-dimensional settings.
method Using variational representations, sample splitting, and self-normalization to create a refined test statistic.
result The resulting statistic has a Gaussian limiting distribution regardless of how dimensionality scales with sample size.

Estimates dimension of subsets from random samples, proving consistency.

problem Estimating the dimension of a compact subset from random samples.
method Consistency proofs for Minkowski, correlation, and pointwise dimensions using empirical volume function.
result Statistical consistency of estimators for various dimension notions.

The paper provides statistical guarantees for generative models using dimension reduction.

problem Improving the quality of generative models without increasing dimensionality.
method Modeling generative devices as smooth transformations of a lower-dimensional space and using integral probability metrics.
result Established a risk bound showing the impact of dimension reduction on generative model error.

Scalability of statistical estimators is of increasing importance in modern applications and dimension reduction is often used to extract relevant information from data. A variety of popular dimension reduction approaches can be framed as symmetric generalized eigendecomposition problems. In this paper we outline how t…

2012-11-07abs ↗pdf ↗

The paper introduces gapped scale-sensitive dimensions to improve learning rate bounds.

problem Improving lower bounds on rates of convergence in statistical and online learning.
method Introducing and analyzing gapped scale-sensitive dimensions for function classes.
result Gapped dimensions lead to stronger lower bounds on offset Rademacher averages.

A method for clustering small datasets in high dimensions using random projections.

problem Challenges in clustering small datasets in high-dimensional spaces.
method Random projection followed by binary clustering in one-dimensional space.
result Statistically significant clustering structures can be found with as few as 100-200 points.

The paper analyzes constrained optimal portfolios in high dimensions using novel statistical learning techniques.

problem Forming optimal portfolios with constraints in high-dimensional asset spaces.
method CROWN method integrating factor models with nodewise regression for estimation in large dimensions.
result Demonstrates estimation consistency and convergence rates for constrained portfolio weights, risk, and Sharpe Ratio.

The scalability of statistical estimators is of increasing importance in modern applications. One approach to implementing scalable algorithms is to compress data into a low dimensional latent space using dimension reduction methods. In this paper we develop an approach for dimension reduction that exploits the assumpt…

2015-04-13abs ↗pdf ↗

The study maps ML quality dimensions to fairness, enhancing the QF4SA framework.

problem Ensuring fairness in ML applications at NSOs to avoid social impacts.
method Employing the QF4SA framework, the study maps quality dimensions to fairness and investigates their interactions.
result Fairness is identified as a new quality dimension in the QF4SA framework.

Study on the Kodaira dimension of real parallelizable manifolds with almost complex structures.

problem Understanding the Kodaira dimension of real parallelizable manifolds with specific almost complex structures.
method Conditions and examples provided for calculating the Kodaira dimension of manifolds.
result Conditions under which the Kodaira dimension of a real parallelizable manifold is zero.

We develop the necessary theory in computational algebraic geometry to place Bayesian networks into the realm of algebraic statistics. We present an algebra{statistics dictionary focused on statistical modeling. In particular, we link the notion of effiective dimension of a Bayesian network with the notion of algebraic…

2012-07-11abs ↗pdf ↗

High-dimensional U-statistics show surprising phase transitions, impacting kernel-based tests.

problem Understanding phase transitions in high-dimensional U-statistics.
method Proved a convergence theorem for U-statistics of degree two in high dimensions.
result High-dimensional U-statistics can have non-Gaussian limits with larger variance and asymmetry.

Develops a computationally tractable high-dimensional differential privacy estimator.

problem Differential privacy in high dimensions is computationally intractable.
method Combines high-dimensional robust statistics with differential privacy techniques.
result A computationally tractable algorithm with dimension-independent privacy loss.

Optimizes ICA performance in high dimensions with computational constraints.

problem Statistical optimality and computational tractability in ICA.
method Characterization of optimal sample complexity, development of computationally tractable estimates.
result Optimal sample complexity is linear in dimensionality, quadratic with low-degree polynomial algorithms.

Unified model for interactive estimation with improved learnability measure.

problem Improving learnability in interactive estimation models.
method Introducing a combinatorial measure (dissimilarity dimension) and a general algorithm with polynomial bounds.
result Unified model subsumes statistical-query learning and structured bandits.

Reduces IB problem to a simpler, lower-dimensional problem.

problem Information bottleneck problem in high-dimensional spaces.
method Identifies sufficient statistic that factors conditional distribution, reducing IB to a lower-dimensional problem.
result Preserves full IB curve and optimal representations, making IB tractable.

For data living in a manifold MRmM\subseteq \mathbb{R}^m and a point pMp\in M we consider a statistic Uk,nU_{k,n} which estimates the variance of the angle between pairs of vectors XipX_i-p and XjpX_j-p, for data points XiX_i, XjX_j, near pp, and evaluate this statistic as a tool for estimation of the intrinsic dimension o…

2018-05-04abs ↗pdf ↗

Characterizes optimal reconstruction error in high-dimensional Gaussian mixtures.

problem Optimizing reconstruction error in high-dimensional sparse Gaussian mixtures.
method Exact asymptotic characterization using state evolution of AMP algorithm.
result Identification of statistical-to-computational gap between AMP and information-theoretic threshold.

Characterizes statistical complexity of realizable regression in PAC and online learning.

problem Understanding the statistical complexity of realizable regression in both PAC and online learning settings.
method Introduces minimax instance optimal learners, novel and combinatorial dimensions to characterize learnability.
result Characterizes which classes of real-valued predictors are learnable and provides necessary conditions for learnability.

Dimension reduction is the process of embedding high-dimensional data into a lower dimensional space to facilitate its analysis. In the Euclidean setting, one fundamental technique for dimension reduction is to apply a random linear map to the data. This dimension reduction procedure succeeds when it preserves certain …

2015-11-30abs ↗pdf ↗

The paper improves GANs' theoretical guarantees for low-dimensional data.

problem Theoretical guarantees for GANs' statistical accuracy remain pessimistic.
method Analytical derivation of statistical guarantees on estimated densities.
result Theoretical rates of convergence for GANs and BiGANs are derived.

Efficient streaming algorithms for robust statistics with near-optimal memory.

problem High-dimensional robust statistics tasks in streaming model.
method First efficient streaming algorithms with near-optimal memory requirements.
result Near-optimal error guarantees and space complexity nearly-linear in the dimension for robust mean estimation.

New smoothing technique improves Wasserstein distance estimation in high dimensions.

problem Estimating statistical distances between high-dimensional distributions.
method Gaussian smoothing of pp-Wasserstein distance and analysis of its asymptotic behavior.
result Gaussian-smoothed pp-Wasserstein distance converges at rate n1/2n^{-1/2}, improving over n1/dn^{-1/d} for unsmoothed distances.

Learning to control linear systems is statistically hard, especially for underactuated systems.

problem Statistical difficulty of learning to control linear systems, especially underactuated ones.
method Utilized minimax lower bounds and structural assumptions to prove learning complexity can be exponential.
result Learning complexity can be at most exponential with the controllability index of the system.

Modified relative universality for unbiasedness and consistency in dimension reduction.

problem Gap in proof of unbiasedness and Fisher consistency in relative universality.
method Modified definition of relative universality using ǫ-measurability.
result Established unbiasedness and Fisher consistency rigorously.

This work is an analytical and numerical study of the composition of several fractals into one and of the relation between the composite dimension and the dimensions of the component fractals. In the case of composition of standard IFS with segments of equal size, the composite dimension can be expressed as a function …

2014-07-10abs ↗pdf ↗

The paper examines extreme value statistics of high-dimensional sample covariances, with applications in finance and image analysis.

problem Statistical validation of normal conditions in high-dimensional time series data.
method Generalizes the maximal deviation of sample autocovariances to high dimensions and applies Gumbel-type extreme value asymptotics.
result Gumbel-type extreme value asymptotics holds true for high-dimensional sample covariances.

Stochastic gradient descent converges to universal limits in high dimensions.

problem Statistical tasks in high dimensions with specific data projections.
method Stochastic gradient descent applied to mixture distributions, proving universality of limits.
result The ODE limits are universal for mixtures of arbitrary product distributions.