Paper combines regularization and pruning to reduce FLOPs in DNNs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We theoretically investigate the convergence rate and support consistency (i.e., correctly identifying the subset of non-zero coefficients in the large sample limit) of multiple kernel learning (MKL). We focus on MKL with block-l1 regularization (inducing sparse kernel combination), block-l2 regularization (inducing un…
Combining explicit and implicit regularization improves deep learning performance without needing depth.
We extend the validity of the Penrose singularity theorem to spacetime metrics of regularity . The proof is based on regularisation techniques, combined with recent results in low regularity causality theory.
Unified parallel ADMM for high-dimensional regression with combined regularizations.
In this work we study input gradient regularization of deep neural networks, and demonstrate that such regularization leads to generalization proofs and improved adversarial robustness. The proof of generalization does not overcome the curse of dimensionality, but it is independent of the number of layers in the networ…
Two important goals of high-dimensional modeling are prediction and variable selection. In this article, we consider regularization with combined and concave penalties, and study the sampling properties of the global optimum of the suggested method in ultra-high dimensional settings. The -penalty provides th…
Multiple kernel learning (MKL), structured sparsity, and multi-task learning have recently received considerable attention. In this paper, we show how different MKL algorithms can be understood as applications of either regularization on the kernel weights or block-norm-based regularization, which is more common in str…
We demonstrate that almost all non-parametric dimensionality reduction methods can be expressed by a simple procedure: regularized loss minimization plus singular value truncation. By distinguishing the role of the loss and regularizer in such a process, we recover a factored perspective that reveals some gaps in the c…
Multi-label text classification is a popular machine learning task where each document is assigned with multiple relevant labels. This task is challenging due to high dimensional features and correlated labels. Multi-label text classifiers need to be carefully regularized to prevent the severe over-fitting in the high …
ε-Consistent Mixup improves semi-supervised classification accuracy.
GOLFS selects features for clustering by combining global and local information.
Hybrid tensor networks improve machine learning by combining quantum and classical methods.
In this paper, we propose -norm regularized models to seek near-optimal sparse portfolios. These sparse solutions reduce the complexity of portfolio implementation and management. Theoretical results are established to guarantee the sparsity of the second-order KKT points of the -norm regularized models…
Paper studies heat flow for maps on manifolds, avoiding singularities.
Batch Normalization is a commonly used trick to improve the training of deep neural networks. These neural networks use L2 regularization, also called weight decay, ostensibly to prevent overfitting. However, we show that L2 regularization has no regularizing effect when combined with normalization. Instead, regulariza…
This study investigates self-supervised learning with Wasserstein distance on tree structures.
Deep learning method improves myelin water fraction estimation.
A new method combines regularization and generative rehearsal for continual learning.
Regularization effect found in neural feature alignment.
pystacked combines machine learning models for improved predictions.
A new method combines generative models and regularization to prevent forgetting in continual learning.
Given a pseudo-Riemannian metric of regularity on a smooth manifold, we prove that the corresponding exponential map is a bi-Lipschitz homeomorphism locally around any point. We also establish the existence of totally normal neighborhoods in an appropriate sense. The proofs are based on regularization, combin…
In recent years a number of methods have been developed for automatically learning the (sparse) connectivity structure of Markov Random Fields. These methods are mostly based on L1-regularized optimization which has a number of disadvantages such as the inability to assess model uncertainty and expensive crossvalidatio…
In recent years a number of methods have been developed for automatically learning the (sparse) connectivity structure of Markov Random Fields. These methods are mostly based on L1-regularized optimization which has a number of disadvantages such as the inability to assess model uncertainty and expensive cross-validati…
Study iterative regularization for linear models with convex bias, improving robust sparse recovery.
Low regularity spacetimes split into simpler structures.
We introduce Implicit Policy, a general class of expressive policies that can flexibly represent complex action distributions in reinforcement learning, with efficient algorithms to compute entropy regularized policy gradients. We empirically show that, despite its simplicity in implementation, entropy regularization c…
We investigate regularized algorithms combining with projection for least-squares regression problem over a Hilbert space, covering nonparametric regression over a reproducing kernel Hilbert space. We prove convergence results with respect to variants of norms, under a capacity assumption on the hypothesis space and a …
This paper analyzes Stochastic Depth regularization in ResNets.
Improves neural network performance with double regularization.
New approach to portfolio optimization shows entropy regularization is ineffective.
Study proves existence of expanding solutions for multiphase surfaces with regular junctions.
Study improves regularity estimates for harmonic maps into ellipsoids.
We introduce a simple and effective method for regularizing large convolutional neural networks. We replace the conventional deterministic pooling operations with a stochastic procedure, randomly picking the activation within each pooling region according to a multinomial distribution, given by the activities within th…
Bayesian Markowitz portfolio problem shows entropy regularization is ineffective.
We propose a new point of view for regularizing deep neural networks by using the norm of a reproducing kernel Hilbert space (RKHS). Even though this norm cannot be computed, it admits upper and lower approximations leading to various practical strategies. Specifically, this perspective (i) provides a common umbrella f…
We show the regularity of, and derive a-priori estimates for (weakly) harmonic maps from a Riemannian manifold into a Euclidean sphere under the assumption that the image avoids some neighborhood of a half-equator. The proofs combine constructions of strictly convex functions and the regularity theory of quasi-linear e…
Enhances financial data signal-to-noise ratio using auto-encoders and mutual regularization.
The paper tackles safe reinforcement learning with convex regularization.
Gibbs pruning optimizes neural networks by combining physics and regularization.
Observational data usually comes with a multimodal nature, which means that it can be naturally represented by a multi-layer graph whose layers share the same set of vertices (users) with different edges (pairwise relationships). In this paper, we address the problem of combining different layers of the multi-layer gra…
KL regularization helps RL algorithms by implicitly averaging q-values.
Optimal ridge regularization computed iteratively from generative parameters.
Mixup improves model accuracy and calibration through data transformation and random perturbation.
Unified framework for understanding and optimizing training acceleration.
This note characterizes monohedral tilings of regular polygons with up to three tiles.
Proves higher regularity for anisotropic inverse mean curvature flow.