In this note, we derive concentration inequalities for random vectors with subGaussian norm (a generalization of both subGaussian random vectors and norm bounded random vectors), which are tight up to logarithmic factors.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper studies how norms of random vectors are preserved by random projections.
We develop time-uniform confidence spheres for estimating means of random vectors.
Study improves error bounds for sparse regression with heavy-tailed covariates.
Random square-tiled surfaces have normal genus distribution and cover all integer vectors.
The paper introduces a new method for tail bounds of random vectors and matrices.
Stochastic trace estimation with tensor train random vectors
Generalizes randomized SVD for better matrix approximations using Gaussian vectors.
The paper generalizes product inequalities for random vectors and their applications.
In this paper we explore the "vector semantics" problem from the perspective of "almost orthogonal" property of high-dimensional random vectors. We show that this intriguing property can be used to "memorize" random vectors by simply adding them, and we provide an efficient probabilistic solution to the set membership …
Paper analyzes error bounds for learning with vector-valued RF, improving existing analyses.
Paper studies tensor models using random matrix theory.
The study examines lower and upper bounds of Wasserstein distances for affine transformations of random vectors.
Market forecasts converge to true values if some agents are correct.
Randomized algorithm solves vector-valued regression problems with low-rank operators.
We introduce a new functional measure of tail dependence for weakly dependent (asymptotically independent) random vectors, termed weak tail dependence function. The new measure is defined at the level of copulas and we compute it for several copula families such as the Gaussian copula, copulas of a class of Gaussian mi…
Optimal transport is #P-hard when components are independent, even with approximate solutions.
Random matrix ensembles yield uniform distributions on manifolds.
We present a new trace estimator of the matrix whose explicit form is not given but its matrix multiplication to a vector is available. The form of the estimator is similar to the Hutchison stochastic trace estimator, but instead of the random noise vectors in Hutchison estimator, we use small number of probing vectors…
The paper shows vector-valued risk measures ignore dependence structures.
We present extremal constructions connected with the property of simplicial collapsibility. (1) For each , there are collapsible (and shellable) simplicial -complexes with only one free face. Also, there are non-evasive -complexes with only two free faces. (Both results are optimal in all dimensions.) (2…
New tree-structured Markov fields with Poisson marginals for counting variables.
Random sinusoidal features are a popular approach for speeding up kernel-based inference in large datasets. Prior to the inference stage, the approach suggests performing dimensionality reduction by first multiplying each data vector by a random Gaussian matrix, and then computing an element-wise sinusoid. Theoretical …
Reconstructing signature features from randomized vector fields in differential equations.
Let and be two -dimensional elliptical random vectors, we establish an identity for , where fulfilling some regularity conditions. Using this identity we provide a unified derivation of sufficient and necessary conditions for classif…
Dictionaries are collections of vectors used for representations of random vectors in Euclidean spaces. Recent research on optimal dictionaries is focused on constructing dictionaries that offer sparse representations, i.e., -optimal representations. Here we consider the problem of finding optimal dictionaries …
Bayesian approach approximates probability functions of Gaussian mixtures.
Study analyzes perturbations in singular subspaces under random noise.
New method reduces deep learning training costs by approximating vector-jacobian products.
Monotone aggregation of dependent random vectors has an absolutely continuous distribution under certain conditions.
New method estimates Gaussian vector functions more efficiently.
The Kaczmarz algorithm is popular for iteratively solving an overdetermined system of linear equations. The traditional Kaczmarz algorithm can approximate the solution in few sweeps through the equations but a randomized version of the Kaczmarz algorithm was shown to converge exponentially and independent of number of …
In a very high-dimensional vector space, two randomly-chosen vectors are almost orthogonal with high probability. Starting from this observation, we develop a statistical factor model, the random factor model, in which factors are chosen at random based on the random projection method. Randomness of factors has the con…
New ensemble SVM model reduces prediction error without choosing best kernel.
Several important families of computational and statistical results in machine learning and randomized algorithms rely on uniform bounds on quadratic forms of random vectors or matrices. Such results include the Johnson-Lindenstrauss (J-L) Lemma, the Restricted Isometry Property (RIP), randomized sketching algorithms, …
In this note, we consider a fixed vector field on and study the distribution of points which lie on the nodal set (of a random spherical harmonic) where is also tangent. We show that the expected value of the corresponding counting function is asymptotic to the eigenvalue with a leading coefficient that i…
This paper proposes a novel type of random forests called a denoising random forests that are robust against noises contained in test samples. Such noise-corrupted samples cause serious damage to the estimation performances of random forests, since unexpected child nodes are often selected and the leaf nodes that the i…
This paper analyzes the variability of Concept Activation Vectors (CAVs).
New insights into tail behavior of heavy-tailed random vectors and processes.
This paper shows that deep learning (DL) representations of data produced by generative adversarial nets (GANs) are random vectors which fall within the class of so-called \textit{concentrated} random vectors. Further exploiting the fact that Gram matrices, of the type with $X=[x_1,\ldots,x_n]\in \mathbb{R}…
We study covariance matrix estimation for the case of partially observed random vectors, where different samples contain different subsets of vector coordinates. Each observation is the product of the variable of interest with a Bernoulli random variable. We analyze an unbiased covariance estimator under this mod…
Random projections are able to perform dimension reduction efficiently for datasets with nonlinear low-dimensional structures. One well-known example is that random matrices embed sparse vectors into a low-dimensional subspace nearly isometrically, known as the restricted isometric property in compressed sensing. In th…
The least-squares support vector machine is a frequently used kernel method for non-linear regression and classification tasks. Here we discuss several approximation algorithms for the least-squares support vector machine classifier. The proposed methods are based on randomized block kernel matrices, and we show that t…
A new method for approximating softmax and Gaussian kernels with reduced error.
The paper predicts responses on out-of-sample nodes using latent positions on unknown curves.
Gaussian random vectors exhibit the loss of dimension phenomena, which relate to their joint survival tail behaviour. Besides, the fact that the components of such vectors are light-tailed complicates the approximations of various multivariate risk measures significantly. In this contribution we derive precise approxim…
Paper studies binary random projections with controllable sparsity patterns for computational and accuracy advantages.
We study the problem of estimating the mean of a random vector given a sample of independent, identically distributed points. We introduce a new estimator that achieves a purely sub-Gaussian performance under the only condition that the second moment of exists. The estimator is based on a novel concept of a…