The paper studies how more data affects prediction risk in high-dimensional models.
problem The impact of increasing data on prediction risk in high-dimensional models.
method Derives central limit theorem and provides finite-sample distribution and confidence interval for prediction risk.
result Demonstrates 'more data hurt' phenomenon in high-dimensional least squares estimation.
High-dimensional SGD limits show surprising dynamics and phase transitions.
problem Understanding SGD in high dimensions and its scaling limits.
method Proving limit theorems for SGD trajectories in high dimensions, choosing summary statistics, initialization, and step-size.
result Critical scaling regime for step-size, new correction term, and complex diffusive limits.
The paper analyzes SGD in high-dimensional networks, revealing new scaling limits.
problem Understanding SGD dynamics in high-dimensional networks.
method Analyzing the effective dynamics of SGD using recent work on the subject.
result A new correction term emerges at the critical scaling regime, changing the phase diagram.
The paper analyzes Q-learning convergence rates with asynchronous updates.
problem Analyzing convergence rates of asynchronous Q-learning algorithms.
method Derives rates of convergence using high-dimensional central limit theorems.
result Establishes a rate of order up to n−1/6log4(nSA) for hyper-rectangles. High-dimensional U-statistics show surprising phase transitions, impacting kernel-based tests.
problem Understanding phase transitions in high-dimensional U-statistics.
method Proved a convergence theorem for U-statistics of degree two in high dimensions.
result High-dimensional U-statistics can have non-Gaussian limits with larger variance and asymmetry.
New insights into SGD and SGD-M in high dimensions.
problem Understanding and comparing SGD and SGD-M in high-dimensional settings.
method Developed high-dimensional scaling limits for SGD-M and online SGD, examining their dynamics and performance.
result SGD-M amplifies high-dimensional effects, potentially degrading performance compared to online SGD.
New CLT for AIPW estimator in high-dimensional settings.
problem Estimating ATE in high-dimensional covariate scenarios.
method Cross-fitting AIPW estimator with well-specified models.
result Established a new CLT for the scaled cross-fit AIPW.
New CLT for SGD in high-dimensional regression provides online inference.
problem Quantifying uncertainty in SGD for high-dimensional regression.
method Established a high-dimensional CLT for online SGD iterates.
result Developed an online approach for estimating variance in CLT.
This paper proves a central limit theorem for differential privacy in high dimensions.
problem Understanding optimal noise distributions for privacy-accuracy trade-offs in high-dimensional settings.
method Developed a central limit theorem approach to analyze differential privacy mechanisms.
result Gaussian mechanisms achieve the optimal privacy-accuracy trade-off in high dimensions.
New method detects changes in high-dimensional data from small samples.
problem Detecting changes in high-dimensional data with limited samples.
method Angular kernel scan framework for detecting marginal distributional shifts.
result Exact population mean factorization and asymptotically distribution-free test.
We provide a way to infer about existence of topological circularity in high-dimensional data sets in Rd from its projection in R2 obtained through a fast manifold learning map as a function of the high-dimensional dataset X and a particular choice of a positive real σ known as band…
Paper generalizes Gaussian universality and CGMT to dependent data, impacting data augmentation in high-dimensional logistic regression.
problem Limitation of Gaussian universality and CGMT in handling dependent data.
method Generalizes Gaussian universality and CGMT to dependent data (block dependence, m-dependence, mixing). Establishes a novel CGMT framework.
result Gaussian universality holds for high-dimensional logistic regression under various types of dependence.
Two-parameter models can learn high-dimensional targets via gradient flow.
problem Learning high-dimensional targets with limited parameters.
method Gradient flow approach for W<d models. result Two-parameter models can learn targets with arbitrarily high success probability.
The study uncovers the breakdown of Gaussian universality in high-dimensional empirical risk minimization.
problem Understanding the breakdown of Gaussian universality in high-dimensional empirical risk minimization.
method Extending the Convex Gaussian Min-Max Theorem to non-Gaussian settings, deriving asymptotic min-max characterizations, and proving asymptotic equivalence of regularizers.
result The projection of the ERM estimator onto a test covariate approximately follows a Gaussian convolution under certain conditions.
Deep networks learn hierarchical functions more efficiently than shallow ones.
problem Understanding the advantage of deep neural networks over shallow models.
method Analytical study of learning dynamics and generalization performance of deep networks compared to shallow ones.
result Deep networks reduce effective dimensionality, enabling learning with fewer samples.
New method finds minimum in noisy data, useful for model selection.
problem Finding the index of the minimum value in noisy observations.
method Developed an asymptotically normal test statistic integrating cross-validation and differential privacy.
result Achieves a favorable bias-variance trade-off in practical scenarios.
We consider the estimation of integrated covariance (ICV) matrices of high dimensional diffusion processes based on high frequency observations. We start by studying the most commonly used estimator, the realized covariance (RCV) matrix. We show that in the high dimensional case when the dimension p and the observati…
Kolmogorov-Arnold Networks promise scalable performance in high dimensions.
problem Curse of dimensionality in multilayer perceptrons.
method Kolmogorov-Arnold representation theorem and interpolation methods.
result Kolmogorov-Arnold Networks achieve true freedom from the curse of dimensionality.
We consider the problem of estimating E[f(U1,…,Ud)], where (U1,…,Ud) denotes a random vector with uniformly distributed marginals. In general, Latin hypercube sampling (LHS) is a powerful tool for solving this kind of high-dimensional numerical integration problem. In the case of depende…
We study Granger causality testing for high-dimensional time series using regularized regressions. To perform proper inference, we rely on heteroskedasticity and autocorrelation consistent (HAC) estimation of the asymptotic variance and develop the inferential theory in the high-dimensional setting. To recognize the ti…
The paper analyzes Kernel Density Estimation in high dimensions with varying data and dimensionality.
problem High-dimensional Kernel Density Estimation with growing data and dimensionality.
method Examines the behavior of Kernel Density Estimators in the regime where both data points and dimensionality grow with a fixed ratio.
result Three distinct statistical regimes are identified for Kernel-based density estimates, each with different statistical properties.
Method upgrades limit theorems to mixing limit theorems for dynamical systems.
problem Improving limit theorems for dynamical systems.
method General method for upgrading limit theorems to mixing limit theorems.
result Mixing limit theorems for specific subbundles of the Kontsevich-Zorich cocycle.
The paper analyzes PLS-SVD in high-dimensional data integration, revealing its strengths and limitations.
problem Understanding the behavior of PLS-SVD in high-dimensional data integration.
method Analysis using random matrix theory and singular value decomposition.
result PLS-SVD exhibits counter-intuitive or limiting behavior in certain regimes and outperforms PCA when detecting common latent subspace.
We consider the problem of uncertainty assessment for low dimensional components in high dimensional models. Specifically, we propose a decorrelated score function to handle the impact of high dimensional nuisance parameters. We consider both hypothesis tests and confidence regions for generic penalized M-estimators. U…
The paper solves optimal bounds for separating data points in high dimensions.
problem Correcting AI errors and analyzing vulnerabilities in high-dimensional data.
method General stochastic separation theorems with optimal probability estimates.
result Explicit and optimal estimates of separation probabilities for important classes of distributions.
Overview of geometric analysis for manifold learning.
problem Analyzing high-dimensional data via spectral embeddings.
method Heat kernel and eigenfunctions on Riemannian manifolds.
result Uniform control of spectral embeddings on key classes of manifolds.
In this note we would like to present "an analysts' point of view" on the Nash-Kuiper theorem and in particular highlight the very close connection to some aspects of turbulence -- a paradigm example of a high-dimensional phenomenon.
Given the observation of a high-dimensional Ornstein-Uhlenbeck (OU) process in continuous time, we proceed to the inference of the drift parameter under a row-sparsity assumption. Towards that aim, we consider the negative log-likelihood of the process, penalized by an ℓ1-penalization (Lasso and Adaptive Lasso). …
Study improves Hayashi-Yoshida estimator for high-dimensional stock covolatility.
problem Inconsistent performance of Hayashi-Yoshida estimator in high dimensions.
method Analyzed the limiting spectral distribution of the Hayashi-Yoshida estimator.
result Established the connection between the estimator's spectrum and the true covariance matrix in high dimensions.
Derives scaling limits and fluctuations for SGD in high dimensions.
problem Understanding SGD behavior in high-dimensional settings with varying noise levels.
method Interacting particle system approach, treating SGD iterates as such, with covariance structure considered.
result Precise three-step phase transition observed in SGD behavior: ballistic, diffusive, then random.
Paper proves CLTs for Q-learning with asynchronous updates.
problem Establishing convergence rates for Q-learning algorithms.
method Polyak-Ruppert averaging, non-asymptotic and functional CLTs.
result Convergence rates in Wasserstein distance for Q-learning.
We rigorously prove statistical physics predictions for non-convex GLMs in high dimensions.
problem Analyzing high-dimensional optimization problems in non-convex Generalized Linear Models.
method Developed a systematic framework using the Gaussian Min-Max Theorem and AMP to rigorously prove replica-symmetric formulas.
result Validated statistical physics predictions for non-convex GLMs, aligning with physicist's conjectures.
New bounds for SGD in high dimensions improve inference efficiency.
problem Quantifying uncertainty in high-dimensional SGD.
method Established non-asymptotic Berry--Esseen bounds for online least-squares SGD.
result Gaussian Central Limit Theorem holds for t≳d1+δ, extending dimensional scaling. The study extends convergence theorems for Ricci-limit spaces with bounded curvature.
problem Understanding convergence properties of Ricci-limit spaces with bounded curvature.
method Establishing C1,α-regularities and applying Fukaya's fibration theorem. result Optimal generalization of Fukaya's fibration theorem to C1,α limit spaces. Paper proves embedding theorem for conformally compact manifolds.
problem Embedding conformally compact manifolds into hyperbolic spaces.
method Proves analogous Nash Embedding Theorem for conformally compact manifolds.
result Conformally compact manifolds can be isometrically embedded into hyperbolic spaces.
New method estimates shape distance in neural representations with limited data.
problem Measuring geometric similarity between high-dimensional network representations.
method Method-of-moments estimator with tunable bias-variance tradeoff.
result New estimator achieves lower bias than standard methods in high-dimensional settings.
Study on sensor fusion algorithms under high dimensional noise.
problem Behavior of sensor fusion algorithms under high dimensional noise.
method Analysis of NCCA and AD algorithms using Gaussian kernel.
result Robustness of NCCA and AD to high dimensional noise depends on SNR and bandwidth selection.
Analyzes SGD dynamics in two-layer networks, bridging different regimes.
problem Understanding SGD dynamics in high-dimensional and mean-field settings.
method Rigorous analysis via deterministic low-dimensional description of sufficient statistics.
result Infinite-width dynamics remains close to a low-dimensional subspace.
A scalable method for accurate inference of low-dimensional parameters in high-dimensional linear regression.
problem Statistical inference for low-dimensional parameters in high-dimensional linear regression models.
method Mean-field variational Bayes approach, focusing on nuisance parameters and conditional distributions.
result Competitive numerical performance and theoretical guarantees for estimation and uncertainty quantification.
To model modern large-scale datasets, we need efficient algorithms to infer a set of P unknown model parameters from N noisy measurements. What are fundamental limits on the accuracy of parameter inference, given finite signal-to-noise ratios, limited measurements, prior information, and computational tractability …
Proves limit curve theorem for incomplete metric spaces, applies to null distance in Lorentzian manifolds.
problem Control of Lorentzian lengths of limit curves in incomplete metric spaces.
method Proves limit curve theorem for incomplete metric spaces and applies to null distance.
result Strong control on Lorentzian lengths of limit curves in Sormani and Vegas' null distance.
We quantify uncertainty in Oja's algorithm's leading eigenvector estimation.
problem Estimating the error of Oja's algorithm's leading eigenvector from streaming data.
method Combining U-statistics, high-dimensional central limit theorems, and multiplier bootstrap.
result Established a weighted χ² approximation for the error between the eigenvector and algorithm output.
Study examines influence diagnostics in high-dimensional M-estimation.
problem Understanding influence diagnostics in high-dimensional settings.
method Characterized the distribution of leave-one-out influences in high-dimensional Gaussian M-estimation.
result The distribution of influences converges to a limiting measure in high-dimensional settings.
Proves CLT for Brownian paths on pinched negative curvature manifolds.
problem Distribution of Brownian paths on pinched negative curvature manifolds.
method Proof of central limit theorem for distances and Green functions.
result Central limit theorem holds for Brownian paths in pinched negative curvature.
Central limit theorem for Green metrics on hyperbolic groups.
problem Proving a central limit theorem for Green metrics on hyperbolic groups.
method Proving a central limit theorem for Green metrics on hyperbolic groups using probability measures and ordering elements.
result Proved a central limit theorem for Green metrics on hyperbolic groups.
Stochastic gradient descent's long-term fluctuations are described by a diffusion limit.
problem Long-term behavior of stochastic gradient descent in non-smooth settings.
method Functional central limit theorem applied to rescaled trajectory of SGD.
result Characterization of long-term fluctuations around the minimizer.
Testing independence is of significant interest in many important areas of large-scale inference. Using extreme-value form statistics to test against sparse alternatives and using quadratic form statistics to test against dense alternatives are two important testing procedures for high-dimensional independence. However…
Study on kernel tests for high-dimensional data, focusing on MMD and CLT.
problem Asymptotic behavior of kernel two-sample tests in high dimensions and large samples.
method Maximum mean discrepancy (MMD) with isotropic kernels, deriving asymptotic expansions and CLT.
result Interplay between moment discrepancy and dimension-and-sample orders in kernel tests.