Developed a new thresholding method that connects soft and hard thresholding.
problem Connecting soft and hard thresholding methods in data analysis.
method Scaled soft thresholding method with empirical scaling values.
result Found two sources of over-fitting in the scaled soft thresholding method.
Study evaluates thresholds for removing noise from DNN weights using random matrix theory.
problem Removing noise from deep neural network weights for better approximation.
method Model weights as signal + noise, use random matrix theory to estimate thresholds, evaluate using cosine similarity.
result Proposed threshold estimation method improves approximation quality.
Sparse reconstruction approaches using the re-weighted l1-penalty have been shown, both empirically and theoretically, to provide a significant improvement in recovering sparse signals in comparison to the l1-relaxation. However, numerical optimization of such penalties involves solving problems with l1-norms in the ob…
Study how firm liquidation regimes affect shareholder value and stability.
problem Balancing shareholder value and financial stability during firm liquidation.
method Modelled forced liquidation in reduced form, solved singular stochastic control problem.
result Combining distress regions below and above ruin threshold improves both shareholder value and firm survival.
PCR robust to noisy, missing, and mixed-valued covariates.
problem Handling noisy, missing, and mixed-valued covariates in PCR.
method PCR is equivalent to HSVT pre-processing; establishes robustness and finite-sample analysis.
result PCR robust to noise, equivalent to RSC, and can learn good predictive models.
STAT-SVD method reduces high-dimensional data sparsity, achieving optimal estimation.
problem Sparse tensor singular value decomposition for high-dimensional data.
method STAT-SVD method with double projection & thresholding scheme.
result STAT-SVD provides sharp thresholding criterion and minimax rate-optimal estimation.
Optimal rank-adaptive matrix estimation from linear measurements.
problem Estimating high-dimensional matrices from linear measurements with adaptive rank selection.
method Combines Least-Squares estimator with universal singular value thresholding.
result Algorithm performance nearly matches fundamental limits.
Recently, a novel family of biologically plausible online algorithms for reducing the dimensionality of streaming data has been derived from the similarity matching principle. In these algorithms, the number of output dimensions can be determined adaptively by thresholding the singular values of the input data matrix. …
Optimal iterative thresholding algorithms improve upon hard and soft thresholding.
problem Optimizing sparsity or rank constraints in optimization problems.
method Developed the notion of relative concavity for thresholding operators, finding a new class of operators that are optimal.
result A new class of thresholding operators, including ℓq thresholding and reciprocal thresholding, achieves the strongest convergence guarantee. Improved iterative hard thresholding for faster, sparser solutions.
problem Finding sparser solutions without sacrificing runtime.
method Adaptive regularization framework applied to iterative hard thresholding.
result Returns solutions with sparsity O(sκ), improving over existing methods. The paper studies phase transitions in random matrices and tensor unfolding for detecting signals.
problem Phase transitions in singular values and vectors of large random matrices.
method Analysis of singular values and vectors of long rectangular random matrices, and tensor unfolding algorithm for asymmetric rank-one spiked tensor models.
result An exact threshold for tensor unfolding to detect signals, independent of unfolding procedure.
Hard thresholding remains efficient for DNN pruning, but smart pruning offers faster accuracy recovery.
problem Efficiently pruning deep neural networks while minimizing accuracy loss.
method Proposes a novel smart pruning algorithm based on difference of convex functions optimization.
result Smart pruning is often orders of magnitude faster than competing approaches while achieving low accuracy degradation.
New RGraSP framework for efficient non-convex optimization.
problem Large-scale non-convex sparsity-constrained optimization problems.
method Relaxed gradient support pursuit with semi-stochastic gradient hard thresholding.
result Our algorithms converge faster with lower per-iteration cost.
Paper develops algorithms to maximize AUC in imbalanced classification.
problem Maximizing AUC in imbalanced classification problems.
method Developed stochastic hard thresholding algorithms to reformulate U-statistics as ERM.
result Proposed algorithm achieves linear convergence rate.
Noise makes learning linear thresholds hard, but algorithms can still learn near-optimal thresholds.
problem Learning linear thresholds in noisy data.
method Exploiting natural assumptions on data-generating process.
result Efficient learning of near-optimal linear thresholds is still possible with small data even in the presence of noise.
ARHT algorithm improves sparsity guarantees in convex optimization.
problem Optimizing convex functions with sparsity constraints.
method Adaptively Regularized Hard Thresholding (ARHT) algorithm.
result ARHT achieves sparsity bound of γ=O(κ), matching theoretical limits.
Unified formula for training dynamics of linear networks combining lazy and balanced regimes.
problem Training dynamics of linear networks in two distinct setups: lazy and balanced/active.
method Unified formula for the evolution of the learned matrix, combining lazy and balanced regimes.
result Unified formula allows for rapid convergence and low rank bias, proving a complete phase diagram.
The truncated singular value decomposition (SVD) of the measurement matrix is the optimal solution to the_representation_ problem of how to best approximate a noisy measurement matrix using a low-rank matrix. Here, we consider the (unobservable)_denoising_ problem of how to best approximate a low-rank signal matrix bur…
AIHT improves online high-dimensional quantile regression by separating support discovery and refinement.
problem Online high-dimensional quantile regression with structural sparsity.
method Adaptive Iterative Hard Thresholding (AIHT) alternates stochastic updates with adaptive hard-thresholding steps.
result AIHT achieves logarithmic regret for the sliding-window objective in high-dimensional settings.
A fast method estimates Gaussian mixture components without iterative fitting.
problem Estimating the number of components in high-dimensional Gaussian mixtures.
method Center data, compute singular values, and count above a threshold.
result The estimator consistently recovers the true number of components under mild separation condition.
This work interprets GELU and related activations via a first-order loss function.
problem Understanding and optimizing activation functions in neural networks.
method Complementary interpretation using the Gaussian first-order loss function.
result Calibrated or learned uniform-threshold gates are competitive and often outperform GELU, ReLU, and SiLU/Swish.
The use of M-estimators in generalized linear regression models in high dimensional settings requires risk minimization with hard L0 constraints. Of the known methods, the class of projected gradient descent (also known as iterative hard thresholding (IHT)) methods is known to offer the fastest and most scalable sol…
New method for high-dimensional manifold-based inference tackles latent responses.
problem Inference on latent right factor vectors in multi-task learning with large numbers of responses and features.
method SOFARI-R method with two variants: one for strongly orthogonal factors and another for weakly orthogonal factors.
result Bias-corrected estimators for latent right factor vectors with asymptotically normal distributions and justified asymptotic variance estimates.
Optimal intervention in economic networks modeled as influence maximization, with hard computational problems.
problem Optimal intervention in economic networks modeled as influence maximization.
method Transformed into influence maximization-like form, with theoretical and practical implications.
result Optimal intervention is NP-hard and cannot be approximated to a constant factor in polynomial time.
IHT improves sparse distribution learning.
problem Learning sparse discrete distributions.
method Iterative hard thresholding as a solution, with a greedy approximate projection.
result IHT achieves state of the art results for sparse distribution learning.
This paper describes a fast algorithm for recovering low-rank matrices from their linear measurements contaminated with Poisson noise: the Poisson noise Maximum Likelihood Singular Value thresholding (PMLSV) algorithm. We propose a convex optimization formulation with a cost function consisting of the sum of a likeliho…
New method for robust regression with near-optimal performance even with high corruption rates.
problem Robust linear regression with response variable corruptions.
method Adaptive hard thresholding for consistent estimation.
result Near-optimal consistent estimation of the true regression vector with 1−o(1) fraction of corruptions. Hard Thresholding Pursuit (HTP) is an iterative greedy selection procedure for finding sparse solutions of underdetermined linear systems. This method has been shown to have strong theoretical guarantee and impressive numerical performance. In this paper, we generalize HTP from compressive sensing to a generic problem …
IntHT solves sparse quadratic regression in sub-quadratic time and space.
problem Sparse quadratic regression in high-dimensional problems.
method Interaction Hard Thresholding (IntHT) is a variant of Iterative Hard Thresholding tailored for quadratic structures.
result IntHT provably converges to a consistent estimate under high-dimensional sparse recovery assumptions.
Randomized SVD shows phase transitions in noisy data.
problem Noise sensitivity of randomized SVD in large rank matrices.
method Analyzed R-SVD under low-rank signal plus noise model.
result R-SVD exhibits BBP-like phase transition with outliers above detectability threshold.
We analyze the local Rademacher complexity of empirical risk minimization (ERM)-based multi-label learning algorithms, and in doing so propose a new algorithm for multi-label learning. Rather than using the trace norm to regularize the multi-label predictor, we instead minimize the tail sum of the singular values of th…
New algorithm resists contamination in high-dimensional regression with optimal performance.
problem Adversarial and measurement errors in high-dimensional data.
method Adversarial Contamination-resistant Iterative Hard Thresholding (AC-IHT) algorithm.
result Achieves minimax near-optimal estimation and signal-adaptive support recovery.
This paper is concerned with the hard thresholding operator which sets all but the k largest absolute elements of a vector to zero. We establish a {\em tight} bound to quantitatively characterize the deviation of the thresholded solution from a given signal. Our theoretical result is universal in the sense that it ho…
Guarantees sparse recovery for neural networks with iterative hard thresholding.
problem Recovering sparse network weights in neural networks.
method Structural properties of sparse network weights and iterative hard thresholding algorithm.
result Simple iterative hard thresholding algorithm recovers sparse network weights exactly using linear memory.
Matrix estimation improves individual fairness without sacrificing performance.
problem Ensuring fairness in algorithmic decision-making.
method Using singular value thresholding (SVT) to preprocess data.
result SVT pre-processing improves IF guarantees and maintains performance.
Variable selection in linear models plays a pivotal role in modern statistics. Hard-thresholding methods such as l0 regularization are theoretically ideal but computationally infeasible. In this paper, we propose a new approach, called the LAGS, short for "least absulute gradient selector", to this challenging yet i…
New findings show good representations alone are insufficient for efficient reinforcement learning.
problem Understanding when good representations are enough for efficient reinforcement learning.
method Statistical analysis of reinforcement learning methods, focusing on value-based, model-based, and policy-based learning.
result Hard thresholds for reinforcement learning methods show good representations alone are insufficient, unless they meet certain quality criteria.
New framework improves GAN training by controlling weight spectra.
problem Training GANs is challenging due to instability and poor generalization.
method Proposes a new reparameterization approach for the discriminator's weight matrices to control spectra.
result Spectrum control enhances GANs' generalization ability and image quality.
Improved IHT with momentum accelerates convex optimization with non-convex constraints.
problem Optimizing convex criteria with non-convex constraints.
method Modified iterative hard thresholding with momentum.
result Acceleration leads to significant improvements over state-of-the-art methods.
Study examines local extrema and crossing statistics in financial markets.
problem Understanding local extrema and crossing statistics in financial markets.
method Excursion set theory, numerical computation, theoretical prediction, clustering of geometrical measures, cross-correlation, Singular Value Decomposition.
result Excursion sets reveal statistical coherency and sensitivity to crises in financial markets.
Proposes a new method for logistic PCA to avoid overfitting.
problem Overfitting in logistic PCA for binary data.
method Non-convex singular value thresholding for logistic PCA.
result Proposed method outperforms models with convex penalties.
Hard phase in inference problems is glassy and hard to reconstruct.
problem Hard phase in inference problems that are hard to solve algorithmically.
method Study of metastable states and their entropy in low-rank matrix factorization.
result AMP algorithm performance is not improved by considering glassy states.
New algorithm robustly estimates sparse models in high dimensions with corrupted data.
problem Estimating latent variable models with arbitrarily corrupted samples in high dimensional space.
method Trimmed (Gradient) Expectation Maximization with trimming gradients and hard thresholding steps.
result The algorithm converges to near optimal statistical rate geometrically under certain conditions.
Paper extends tensor recovery method for low CP-rank tensors.
problem Recovery of low-rank tensors from few measurements.
method Iterative Hard Thresholding with tensor version of RIP.
result Exact recovery of tensors with low CP-rank is guaranteed.
This work analyzes self-attention matrices using random matrix theory.
problem Understanding the theoretical behavior of self-attention layers in neural networks.
method Asymptotic spectral analysis of the attention matrix, Gaussian equivalence, and linearization.
result The singular value distribution of the attention matrix is asymptotically characterized by a linear model.
Training neural networks is hard in fixed dimensions.
problem Training two-layer neural networks is computationally hard in fixed dimensions.
method Parameterized complexity analysis considering dimension and number of neurons.
result Training two-layer neural networks is NP-hard for two dimensions.
Learning β for k-SAT with one sample is hard, especially for low degrees.
problem Learning the parameter β for k-SAT with a single sample.
method Single-sample learning of k-SAT formulas with a given satisfying assignment.
result One-shot learning for k-SAT is infeasible well below the satisfiability threshold, even for low degrees.
Several learning applications require solving high-dimensional regression problems where the relevant features belong to a small number of (overlapping) groups. For very large datasets and under standard sparsity constraints, hard thresholding methods have proven to be extremely efficient, but such methods require NP h…