In this paper we consider regularized convex cone programming problems. In particular, we first propose an iterative hard thresholding (IHT) method and its variant for solving regularized box constrained convex programming. We show that the sequence generated by these methods converges to a local minimizer.…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We study the estimation of for the nonlinear model $y = f(X\sp{\top}β) + ε$ when is a nonlinear transformation that is known, has sparse nonzero coordinates, and the number of observations can be much smaller than that of parameters (). We show that in order to bound the error of the reg…
NACT improves tensor regression predictions with regularization.
We consider a discrete optimization formulation for learning sparse classifiers, where the outcome depends upon a linear combination of a small subset of features. Recent work has shown that mixed integer programming (MIP) can be used to solve (to optimality) -regularized regression problems at scales much larg…
We develop a primal dual active set with continuation algorithm for solving the \ell^0-regularized least-squares problem that frequently arises in compressed sensing. The algorithm couples the the primal dual active set method with a continuation strategy on the regularization parameter. At each inner iteration, it fir…
Bayesian -regularized least squares is a variable selection technique for high dimensional predictors. The challenge is optimizing a non-convex objective function via search over model space consisting of all possible predictor combinations. Spike-and-slab (a.k.a. Bernoulli-Gaussian) priors are the gold standard f…
Variable selection in linear models plays a pivotal role in modern statistics. Hard-thresholding methods such as regularization are theoretically ideal but computationally infeasible. In this paper, we propose a new approach, called the LAGS, short for "least absulute gradient selector", to this challenging yet i…
Paper proposes a new sparse group k-max regularization for sparsity constraints.
In a recent work (arXiv:0910.2517), for nonlinear models with sparse underlying linear structures, we studied the error bounds of -regularized estimation. In this note, we show that -regularized estimation in some important cases can achieve the same order of error bounds as those in the aforementioned …
Paper tackles low-rank matrix recovery with column -norm regularization.
We propose a practical method for norm regularization for neural networks: pruning the network during training by encouraging weights to become exactly zero. Such regularization is interesting since (1) it can greatly speed up training and inference, and (2) it can improve generalization. AIC and BIC, well-known …
The -regularized least squares problem (a.k.a. best subsets) is central to sparse statistical learning and has attracted significant attention across the wider statistics, machine learning, and optimization communities. Recent work has shown that modern mixed integer optimization (MIO) solvers can be used to addre…
L0Learn solves sparse learning problems with millions of features.
We show that any simply connected topological closed -manifold punctured along any compact, totally disconnected tame subset admits a continuum of smoothings which are not diffeomorphic to any leaf of a codimension one foliation on a compact manifold. This includes the remarkable case of puncture…
Despite its nonconvex nature, sparse approximation is desirable in many theoretical and application cases. We study the sparse approximation problem with the tool of deep learning, by proposing Deep Encoders. Two typical forms, the regularized problem and the -sparse problem, are …
Optimizes subset selection in sparse learning problems.
Efficient algorithm solves best subset selection problem.
Given two data matrices and , sparse canonical correlation analysis (SCCA) is to seek two sparse canonical vectors and to maximize the correlation between and . However, classical and sparse CCA models consider the contribution of all the samples of data matrices and thus cannot identify an unde…
Safe screening rules reduce -regression computation by fixing 76% of variables.
Sparse Bayesian Optimization (SEBO) finds interpretable configurations.
New solver SR2 tackles deep neural network training with nonsmooth regularization.
The search for efficient, sparse deep neural network models is most prominently performed by pruning: training a dense, overparameterized network and removing parameters, usually via following a manually-crafted heuristic. Additionally, the recent Lottery Ticket Hypothesis conjectures that, for a typically-sized neural…
High-dimensional sparse modeling via regularization provides a powerful tool for analyzing large-scale data sets and obtaining meaningful, interpretable models. The use of nonconvex penalty functions shows advantage in selecting important features in high dimensions, but the global optimality of such methods still dema…
A new method prunes neural networks efficiently without losing effectiveness.
Efficiently infers time-varying sparse MRFs with strong statistical guarantees.
In many applications, high-dimensional data points can be well represented by low-dimensional subspaces. To identify the subspaces, it is important to capture a global and local structure of the data which is achieved by imposing low-rank and sparseness constraints on the data representation matrix. In low-rank sparse …
New method selects sparse predictors in large LMMs.
We solve the regularity problem for Milnor's infinite dimensional Lie groups in the asymptotic estimate context. Specifically, let be a Lie group with asymptotic estimate Lie algebra , and denote its evolution map by , i.e.…
Proposes -CCA for sparse CCA with improved representation learning.
LTP learns per-layer thresholds for efficient pruning of deep networks.
Enhanced kernel ridgeless regression improves performance with LAB RBF kernels.
New algorithm solves complex variable selection problems in high dimensions.
Improves convergence speed in compressive sensing with a new probabilistic approach.
Quantization can be used to form new vectors/matrices with shared values close to the original. In recent years, the popularity of scalar quantization for value-sharing applications has been soaring as it has been found huge utilities in reducing the complexity of neural networks. Existing clustering-based quantization…