Proposes flexible auto-encoders for varying data dimensions.
problem Fixed latent dimensions limit data flexibility.
method Stochastic bottleneck with weighted dropouts.
result Seamless variable dimensionality reduction with high performance.
Develops a high-dimensional differentially-private EM algorithm with near-optimal statistical guarantees.
problem Designing differentially-private EM algorithms for high-dimensional latent variable models.
method Noisy iterative hard-thresholding, statistical guarantees, near-optimal convergence rates.
result Near-optimal statistical guarantees and minimax rate optimality in high-dimensional settings.
We propose a novel sparse tensor decomposition method, namely Tensor Truncated Power (TTP) method, that incorporates variable selection into the estimation of decomposition components. The sparsity is achieved via an efficient truncation step embedded in the tensor power iteration. Our method applies to a broad family …
Paper develops a novel approach for classifying high-dimensional mixed data.
problem Handling datasets with both categorical and continuous variables of high dimensions.
method Location model with Gaussian conditional distributions, kernel smoothing for bandwidth choice, penalized likelihood estimation.
result Competitive performance of the proposed classifier demonstrated through simulations and real data.
T-Rex selector selects variables fast and controls FDR in high-dimensional data.
problem Variable selection in high-dimensional data with FDR control.
method Fused solutions of early terminated random experiments.
result FDR control at target level with high variable selection power.
Simplifies IV regression for high-dimensional instruments.
problem Nonlinear instrumental variable regression with high-dimensional instruments.
method Combines kernelized IV methods with an adaptive regression algorithm.
result Faster convergence and adaptability to feature dimensionality.
Enhances FDR control in variable selection using neural networks.
problem Balancing rigorous error control with statistical power in high-dimensional variable selection.
method Learning-augmented T-Rex Selector framework with a neural network trained on synthetic datasets.
result Achieves superior detection of true variables compared to existing approaches.
Study on reducing dimensionality in high-dimensional regression with kernel methods and stability analysis.
problem Analyzing errors in high-dimensional regression with dimensionality reduction and kernel regression.
method Derive a stability result for kernel regression with Wasserstein distance and apply it to PCA to deduce convergence rates.
result Two-step procedure yields useful convergence rates in semi-supervised settings.
In their activity, the traders approximate the rate of return by integer multiples of a minimal one. Therefore, it can be regarded as a quantized variable. On the other hand, there is the impossibility of observing the rate of return and its instantaneous forward time derivative, even if we consider it as a continuous …
Paper proposes knockoff-based methods to simplify deep neural networks by controlling false discovery rates.
problem High-dimensional deep neural networks with many irrelevant parameters and inputs.
method Knockoff methods combined with regularized neural networks for variable screening.
result Proposed algorithms show satisfactory performance in controlling false discovery rates.
ISOKANN learns collective variables and effective dynamics for metastable transitions.
problem Understanding metastable transitions in complex molecular systems.
method Integrates Koopman operators with neural networks to extract CVs and effective dynamics.
result Reconstructs coarse-grained kinetics and reproduces transition times across barriers.
Novel framework controls FDR in high-dimensional, dependent data.
problem FDR control failure in high-dimensional, dependent data.
method Dependency-aware T-Rex selector integrating hierarchical graphical models and martingale theory.
result First to control FDR in high-dimensional, dependent data.
We consider the high-dimensional discriminant analysis problem. For this problem, different methods have been proposed and justified by establishing exact convergence rates for the classification risk, as well as the l2 convergence results to the discriminative rule. However, sharp theoretical analysis for the variable…
Sparse PCA selects variables with FDR control for improved performance.
problem Sparse PCA selects irrelevant variables when maximizing explained variance.
method Proposes FDR-controlled selection using T-Rex selector.
result Significant performance improvement over traditional sparse PCA.
This paper explores the following question: what kind of statistical guarantees can be given when doing variable selection in high-dimensional models? In particular, we look at the error rates and power of some multi-stage regression methods. In the first stage we fit a set of candidate models. In the second stage we s…
Bayesian approach controls FDR in high-dimensional models.
problem High-dimensional variable selection and inference.
method Adapted Mirror Statistic to Bayesian framework for FDR control.
result Effective FDR control without data splitting.
Paper improves deep learning convergence rates for low-dimensional data.
problem Sub-optimal rates in deep learning due to unrealistic assumptions on intrinsic dimension.
method Introduced an entropic notion of intrinsic dimension for exponential families and demonstrated improved convergence rates.
result Test error scales as O~(n−2β+dˉ2β(λ)2β), improving on best-known rates. Novel privatization framework for high-dimensional variable selection with differential privacy.
problem High-dimensional controlled variable selection with rigorous FDR control under differential privacy constraints.
method Gaussian Johnson-Lindenstrauss Transformation for privatizing the knockoff matrix.
result The proposed private variable selection procedure maintains statistical power even under strict privacy budgets.
We present a finite-dimensional version of the quantum model for the stock market proposed in [C. Zhang and L. Huang, A quantum model for the stock market, Physica A 389(2010) 5769]. Our approach is an attempt to make this model consistent with the discrete nature of the stock price and is based on the mathematical for…
Improves change-point detection for high-dimensional time-series.
problem Uncertainty in latent variable estimation affects change-point detection.
method Proposes multinomial sampling to improve detection rate and reduce delay.
result Results outperform baseline method in experiments.
This work develops rigorous theoretical basis for the fact that deep Bayesian neural network (BNN) is an effective tool for high-dimensional variable selection with rigorous uncertainty quantification. We develop new Bayesian non-parametric theorems to show that a properly configured deep BNN (1) learns the variable im…
This paper proposes a general adaptive procedure for budget-limited predictor design in high dimensions called two-stage Sampling, Prediction and Adaptive Regression via Correlation Screening (SPARCS). SPARCS can be applied to high dimensional prediction problems in experimental science, medicine, finance, and engineer…
New CVs preserve transition rates in molecular dynamics.
problem Designing CVs that accurately capture rare events in high-dimensional systems.
method Integrating manifold learning and group-invariant featurization to construct neural network-based CVs that satisfy orthogonality conditions.
result Achieved a CV for butane that reproduces the anti-gauche transition rate with less than ten percent relative error.
New method improves simulation efficiency in high dimensions.
problem Efficiency in estimating functionals of conditional expectations in high dimensions.
method Kernel ridge regression exploiting smoothness of conditional expectation.
result Effective reduction of the curse of dimensionality, bridging convergence rates.
Learning rate needs to decrease with higher data moments for effective ICA in high dimensions.
problem Slower convergence of ICA in high-dimensional data with high-order moments.
method High-dimensional ODE analysis of ICA algorithm under controlled moment structure.
result Critical learning rate threshold for effective ICA when moments are high.
Paper proposes a privacy-preserving knockoff inference method.
problem Ensuring privacy in model-X knockoff inference.
method Differential privacy framework for knockoff inference.
result Guaranteed FDR control with privacy protection.
We provide a general theory of the expectation-maximization (EM) algorithm for inferring high dimensional latent variable models. In particular, we make two contributions: (i) For parameter estimation, we propose a novel high dimensional EM algorithm which naturally incorporates sparsity structure into parameter estima…
The paper develops a method for optimal projection selection in high-dimensional classification.
problem High-dimensional classification with latent variable structure.
method Formulates a latent-variable model and proposes a computationally efficient classifier.
result Explicit rates of convergence for excess risk of the proposed classifier are derived and shown to be optimal.
New method uses sparse deep neural networks for high-dimensional regression with improved parameter estimation.
problem Improving parameter estimation in high-dimensional sparse regression models.
method Proposes nonparametric estimation of partial derivatives in sparse deep neural networks.
result Established convergence rate of nonparametric estimation of partial derivatives as O(n−1/4). It is now known that an extended Gaussian process model equipped with rescaling can adapt to different smoothness levels of a function valued parameter in many nonparametric Bayesian analyses, offering a posterior convergence rate that is optimal (up to logarithmic factors) for the smoothness class the true function be…
A variable annuity contract with Guaranteed Minimum Withdrawal Benefit (GMWB) promises to return the entire initial investment through cash withdrawals during the contract plus the remaining account balance at maturity, regardless of the portfolio performance. Under the optimal(dynamic) withdrawal strategy of a policyh…
When can reliable inference be drawn in the "Big Data" context? This paper presents a framework for answering this fundamental question in the context of correlation mining, with implications for general large scale inference. In large scale data applications like genomics, connectomics, and eco-informatics the dataset…
A novel AIRLS algorithm for multiaffine variable relations in high-dimensional problems.
problem Challenges in Maximum Likelihood Estimation in high-dimensional settings with complex variable relations.
method Proposes an Alternating and Iteratively-Reweighted Least Squares (AIRLS) algorithm for multiaffine variable relations.
result Proves convergence for problems with Generalized Normal Distributions and shows empirically super-linear convergence rate.
Paper presents deep LSMC method for efficient variable annuity pricing.
problem Efficiently pricing variable annuities with guarantees using simulation methods.
method Modifies least-squares Monte Carlo (LSMC) algorithm for optimal stochastic control problems.
result Deep LSMC provides more stable and robust pricing performance for higher-dimensional problems.
Assigning significance in high-dimensional regression is challenging. Most computationally efficient selection algorithms cannot guard against inclusion of noise variables. Asymptotically valid p-values are not available. An exception is a recent proposal by Wasserman and Roeder (2008) which splits the data into two pa…
Variable selection in high-dimensional space characterizes many contemporary problems in scientific discovery and decision making. Many frequently-used techniques are based on independence screening; examples include correlation ranking (Fan and Lv, 2008) or feature selection using a two-sample t-test in high-dimension…
Proposes a method for stable variable selection in high-dimensional data.
problem Challenges of variable selection in high-dimensional, correlated data.
method Resample-aggregate framework using diffusion models.
result Stable subset of predictors with calibrated stability scores.
New method clusters matrix-valued data by latent variables.
problem Clustering matrix-valued data with hidden structure.
method Latent variable model with hierarchical clustering.
result Algorithm attains clustering consistency in high dimensions.
This paper analyzes deep federated learning for low-dimensional data, revealing intrinsic dimensionality's role in convergence rates.
problem Insufficient investigation of generalization error in heterogeneous federated learning, especially for low-dimensional data.
method Statistical analysis of deep federated regression in a two-stage sampling model.
result Intrinsic dimensionality, characterized by entropic dimension, determines convergence rates for deep learners.
The paper studies implicit regularization in over-parameterized models for high-dimensional data.
problem Understanding implicit regularization in over-parameterized models for high-dimensional data.
method The paper designs regularization-free algorithms for the high-dimensional single index model and provides theoretical guarantees for the induced implicit regularization phenomenon.
result The proposed methods achieve minimax optimal statistical rates of convergence and outperform classical methods with explicit regularization.
Develops methods to estimate high rank tensors from noisy data.
problem Estimating high rank tensors from noisy observations.
method Generative latent variable tensor model, polynomial-time spectral algorithm.
result Achieves computationally optimal rate for signal tensor estimation.
High-dimensional data analysis has motivated a spectrum of regularization methods for variable selection and sparse modeling, with two popular classes of convex ones and concave ones. A long debate has been on whether one class dominates the other, an important question both in theory and to practitioners. In this pape…
Study improves model estimation and variable selection using GANs with Lasso penalty.
problem Variable selection in high-dimensional data with deep networks.
method Conditional Wasserstein Generative Adversarial Networks with Group Lasso penalization.
result Established convergence rate for variable selection in censored survival data.
Proposes a two-stage method for testing variable interactions with FDR control.
problem Testing pairwise interactions in high-dimensional data with dependence.
method Two-stage testing procedure with FDR control using Cramér type moderate deviation technique.
result The proposed method controls FDR and has comparable or improved statistical power.
New method reduces memory usage for high-dimensional variable selection.
problem Scalability issues in high-dimensional variable selection, especially in genomics.
method Adaptive sampling of null features to eliminate dummy matrix materialization.
result Reduces memory and runtime by several orders of magnitude while preserving FDR control.
New algorithms learn latent variable models without tuning, outperforming existing methods.
problem Learning latent variable models without manual tuning.
method Two particle-based algorithms using free energy minimization and coin betting.
result Learning algorithms are entirely tuning-free and competitive with existing methods.
Big T-Rex solves FDR-controlled sparse regression on laptops with millions of variables.
problem Scalable FDR-controlled variable selection for high-dimensional data.
method Early terminated random experiments with memory-mapping and permutation-based dummy generation.
result Solves FDR-controlled Lasso problems with 5 million variables on a laptop in 30 minutes.
New method for accurate permutation inference in CCA.
problem Inaccurate permutation inference in CCA.
method Proposed solutions for permutation inference in CCA, including transforming residuals and stepwise estimation.
result Valid permutation tests for CCA with and without nuisance variables.