For high dimensional data, some of the standard statistical techniques do not work well. So modification or further development of statistical methods are necessary. In this paper, we explore these modifications. We start with the important problem of estimating high dimensional covariance matrix. Then we explore some …
High-dimensional statistics advances in complex data domains.
problem Complex, rich datasets challenge traditional methods.
method Evolved to address sophisticated estimation and inference problems.
result Deepened connections with optimization, concentration, and information theory.
Enhances power of covariance matrix tests for high-dimensional data.
problem Testing large covariance matrices in high-dimensional data.
method Proposes a new Fisher's combined probability test for quadratic form and maximum form statistics.
result Boosts power against more general alternatives.
Efficient streaming algorithms for robust statistics with near-optimal memory.
problem High-dimensional robust statistics tasks in streaming model.
method First efficient streaming algorithms with near-optimal memory requirements.
result Near-optimal error guarantees and space complexity nearly-linear in the dimension for robust mean estimation.
Overview of high-dimensional time series regression methods.
problem Estimation and inference with high-dimensional time series data.
method Limit theory for high-dimensional dependent data, asymptotic theory for time series regression, statistical learning methods.
result Main limit theory results and asymptotic theory for high-dimensional time series regression.
The paper provides statistical guarantees for SGD and ASGD in high-dimensional settings.
problem Theoretical understanding of SGD and ASGD in high-dimensional settings.
method Transfer of tools from high-dimensional time series to online learning, using coupling techniques.
result Established geometric-moment contraction and q-th moment convergence of SGD and ASGD. The paper provides bounds for high-dimensional U-statistics with novel order-explicit inequalities.
problem Bounding the deviation of high-dimensional U-statistics from their Hájek projections.
method Develops novel order-explicit moment inequalities for higher-order Hoeffding components.
result The maximum deviation of a high-dimensional U-statistic from its Hájek projection is of order Op(φbn−1log2(dn)). Unified tutorial on AMP for high-dimensional problems.
problem Structured high-dimensional statistical problems.
method Statistical perspective of AMP and its applications.
result Unified and strengthened results in AMP literature.
New method for estimating high-dimensional binary time series coefficients.
problem Statistical inference for high-dimensional binary time series.
method Post-selection estimator and second-order wild bootstrap algorithm.
result Good finite-sample performance of the proposed method.
Develops a high-dimensional differentially-private EM algorithm with near-optimal statistical guarantees.
problem Designing differentially-private EM algorithms for high-dimensional latent variable models.
method Noisy iterative hard-thresholding, statistical guarantees, near-optimal convergence rates.
result Near-optimal statistical guarantees and minimax rate optimality in high-dimensional settings.
Proposes a method to compare noisy high-dimensional datasets with low-dimensional manifolds.
problem Comparing distributions on manifolds in noisy high-dimensional datasets.
method Linking low-rank structure to manifold geometry, developing a scale-invariant distance measure.
result Superior robustness and statistical power compared to existing methods.
Factor models are a class of powerful statistical models that have been widely used to deal with dependent measurements that arise frequently from various applications from genomics and neuroscience to economics and finance. As data are collected at an ever-growing scale, statistical machine learning faces some new cha…
Study replicability in high-dimensional statistics, resolving open problems.
problem Ensuring consistent results in high-dimensional statistical tasks.
method Introduced replicable learning algorithms and established computational and statistical equivalence with high-dimensional isoperimetric tilings.
result Matching sample complexity upper and lower bounds for replicable mean estimation and coin problem.
Paper introduces PTL-SI for statistical inference in TL-HDR, controlling FPR.
problem Quantifying statistical significance in TL-HDR with limited data.
method PTL-SI framework for valid p-values in TL-HDR feature selection. result Valid p-values and controlled FPR in TL-HDR feature selection. Develops methods for GWAS of high dimensional phenotypes using summary statistics.
problem Lack of methods to model pleiotropy in multi-phenotype GWAS.
method Bayesian inference model using summary statistics, fast computation, and biologically informed priors.
result Demonstrates utility in metabolite GWAS with interpretable pathway-level inference.
High-dimensional U-statistics show surprising phase transitions, impacting kernel-based tests.
problem Understanding phase transitions in high-dimensional U-statistics.
method Proved a convergence theorem for U-statistics of degree two in high dimensions.
result High-dimensional U-statistics can have non-Gaussian limits with larger variance and asymmetry.
New algorithms improve privacy in statistical estimation by making them robust.
problem Improving privacy in statistical estimation methods.
method Black-box reduction from privacy to robustness, using Sum-of-Squares method.
result Design of polynomial-time private estimators with optimal tradeoffs among sample complexity, accuracy, and privacy.
New statistical inference method for high-dimensional Hawkes processes.
problem Uncertainty evaluation of network estimates in high-dimensional point process data.
method Develops a new statistical inference procedure using concentration inequalities and martingale central limit theory.
result Characterizes the convergence rate of test statistics for high-dimensional Hawkes processes.
Testing independence is of significant interest in many important areas of large-scale inference. Using extreme-value form statistics to test against sparse alternatives and using quadratic form statistics to test against dense alternatives are two important testing procedures for high-dimensional independence. However…
We provide a general theory of the expectation-maximization (EM) algorithm for inferring high dimensional latent variable models. In particular, we make two contributions: (i) For parameter estimation, we propose a novel high dimensional EM algorithm which naturally incorporates sparsity structure into parameter estima…
Bayesian approach controls FDR in high-dimensional models.
problem High-dimensional variable selection and inference.
method Adapted Mirror Statistic to Bayesian framework for FDR control.
result Effective FDR control without data splitting.
Noise Sensitivity Exponent controls statistical-computational gaps in learning.
problem Understanding when learning is statistically possible yet computationally hard in high-dimensional statistics.
method Investigating statistical-computational gaps in single- and multi-index models using Noise Sensitivity Exponent.
result Noise Sensitivity Exponent governs statistical-computational gaps in high-dimensional learning.
This paper develops dimension-agnostic inference methods for high-dimensional data.
problem Understanding how classical inference methods behave in high-dimensional settings.
method Using variational representations, sample splitting, and self-normalization to create a refined test statistic.
result The resulting statistic has a Gaussian limiting distribution regardless of how dimensionality scales with sample size.
The paper reviews and improves concentration inequalities for statistical inference.
problem Analyzing statistical inference in various settings with high-dimensional data.
method Review and improvement of concentration inequalities for different types of random variables and statistical measures.
result Fresh new results and improved bounds with sharper constants.
Statistical query algorithms and low-degree tests are nearly equivalent in high-dimensional hypothesis testing.
problem High-dimensional hypothesis testing and information-computation gaps.
method Analysis of statistical query framework and low-degree polynomials.
result Statistical query algorithms and low-degree polynomials are almost equivalent in power under mild conditions.
Paper develops a distributed debiased estimator for sparse statistical inference.
problem High computational costs in debiased estimator construction for high-dimensional models.
method Develops a multi-round distributed debiased estimator using both labeled and unlabelled data.
result Unlabeled data improves statistical rate of each iteration in distributed setup.
We propose a novel sparse tensor decomposition method, namely Tensor Truncated Power (TTP) method, that incorporates variable selection into the estimation of decomposition components. The sparsity is achieved via an efficient truncation step embedded in the tensor power iteration. Our method applies to a broad family …
Approximate Bayesian Computation is widely used in systems biology for inferring parameters in stochastic gene regulatory network models. Its performance hinges critically on the ability to summarize high-dimensional system responses such as time series into a few informative, low-dimensional summary statistics. The qu…
Learning in the presence of outliers is a fundamental problem in statistics. Until recently, all known efficient unsupervised learning algorithms were very sensitive to outliers in high dimensions. In particular, even for the task of robust mean estimation under natural distributional assumptions, no efficient algorith…
Paper explores differential privacy in high-dimensional federated learning, tackling server trustworthiness and estimation.
problem Maintaining privacy in distributed environments with high-dimensional data.
method Investigates scenarios with untrusted and trusted central servers, introduces novel federated estimation algorithms for linear regression models.
result Tight minimax rates depend on high-dimensionality even with sparsity assumptions, and novel algorithms handle slight variations among distributed models.
Develops a computationally tractable high-dimensional differential privacy estimator.
problem Differential privacy in high dimensions is computationally intractable.
method Combines high-dimensional robust statistics with differential privacy techniques.
result A computationally tractable algorithm with dimension-independent privacy loss.
HI-SIGMA improves sensitivity in high-dimensional statistical inference with data-driven background models.
problem Performing high-dimensional statistical inference with complex backgrounds in high-energy physics.
method HI-SIGMA uses generative ML models to learn signal and background distributions, incorporating systematic uncertainties.
result HI-SIGMA provides improved sensitivity compared to classifier-based methods.
We study high-dimensional Gaussian mixture classification using statistical physics methods.
problem Classifying high-dimensional Gaussian mixture with general covariance matrices.
method Replica method from statistical physics for asymptotic analysis of convex classifiers.
result Construction and validation of a de-biased estimator for variable selection.
Paper connects free-energy and low-degree hardness in high-dimensional statistics.
problem High-dimensional statistical inference problems are computationally hard.
method Defines a free-energy criterion and connects it to low-degree hardness.
result Establishes connection between free-energy and low-degree hardness for Gaussian models.
Cross-sectional "Information Coefficient" (IC) is a widely and deeply accepted measure in portfolio management. The paper gives an insight into IC in view of high-dimensional directional statistics: IC is a linear operator on the components of a centralizing-unitizing standardized random vector of next-period cross-sec…
Sparse Polyak improves high-dimensional statistical estimation.
problem High-dimensional statistical estimation problems with growing problem dimension.
method Sparse Polyak modifies Polyak's adaptive step size to estimate restricted Lipschitz smoothness.
result Sparse Polyak achieves optimal statistical precision with fewer iterations.
Density Estimation is one of the central areas of statistics whose purpose is to estimate the probability density function underlying the observed data. It serves as a building block for many tasks in statistical inference, visualization, and machine learning. Density Estimation is widely adopted in the domain of unsup…
Study examines mean estimation in high dimensions with small data.
problem Efficiently estimating mean in high-dimensional data with limited data size.
method Extensive experimentation of various mean estimation techniques.
result Developed robust methods for mean estimation with low data size.
Simple model explains manifold structure in high-dimensional data.
problem Understanding manifold structure in high-dimensional data.
method Latent Metric Model with latent variables, correlation, and stationarity.
result Establishes statistical explanation for manifold hypothesis.
Many statistical M-estimators are based on convex optimization problems formed by the combination of a data-dependent loss function with a norm-based regularizer. We analyze the convergence rates of projected gradient and composite gradient methods for solving such problems, working within a high-dimensional framewor…
Paper develops a new test for high-dimensional matrix-valued data.
problem Hypothesis testing for mean of matrix-valued data in high-dimensional settings.
method Proposes a new test statistic for high-dimensional matrix rank testing.
result Develops a novel approach for sparse singular value decomposition (SVD) estimation.
Characterizes optimal reconstruction error in high-dimensional Gaussian mixtures.
problem Optimizing reconstruction error in high-dimensional sparse Gaussian mixtures.
method Exact asymptotic characterization using state evolution of AMP algorithm.
result Identification of statistical-to-computational gap between AMP and information-theoretic threshold.
New test for comparing high-dimensional text data.
problem Testing equality of multinomial distributions in high dimensions.
method Proposed a test statistic with asymptotic normality under null.
result Achieves optimal detection boundary across parameter space.
Efficiently transforms samples from various statistical models.
problem Approximately transforming samples from one statistical model to another without knowing the source model's parameters.
method Constructs computationally efficient procedures to reduce uniform, Erlang, and Laplace models to general target families.
result Establishes nonasymptotic reductions between canonical high-dimensional problems, such as mixtures of experts, phase retrieval, and signal denoising.
Paper proposes differentially private quantile regression for high-dimensional data.
problem Privacy concerns in big data with heterogeneous sensitive personal information.
method Newton-type transformation for reformulating quantile regression into an OLS problem; iterative updates for estimation; debiased estimator for inference; communication-efficient bootstrap.
result Near-optimal statistical accuracy and formal privacy guarantees achieved.
Novel tests for genetic independence in high-dimensional data.
problem Testing independence in genetics studies with many variables.
method Defining premetric structures on genetic data support spaces.
result Solid theoretical framework and computationally-efficient implementations.
The present paper provides a study of high-dimensional statistical arbitrage that combines factor models with the tools from stochastic control, obtaining closed-form optimal strategies which are both interpretable and computationally implementable in a high-dimensional setting. Our setup is based on a general statisti…
Two methods monitor high-dimensional processes via manifold fitting or learning.
problem Monitoring high-dimensional, dynamic industrial processes.
method Manifold fitting and learning approaches for online SPC.
result Manifold-fitting approach achieves performance competitive with classical methods.