Sparse activations in neural networks are hard to exploit but lead to advantages in learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Recent literature on unsupervised learning focused on designing structural priors with the aim of learning meaningful features, but without considering the description length of the representations. In this thesis, first we introduce the metric that evaluates unsupervised models based on their reconstruction …
Previous literature on unsupervised learning focused on designing structural priors with the aim of learning meaningful features. However, this was done without considering the description length of the learned representations which is a direct and unbiased measure of the model complexity. In this paper, first we intro…
Novel active learning framework using sparse approximation for efficient model training.
BatchTopK SAEs improve GPT-2 and Gemma activations with adjustable sparsity.
REALITrees uses a Rashomon ensemble approach for active learning in sparse decision trees.
Paper hypothesizes MLP layers in LLMs can be approximated by sparse Mixture of Experts.
Picasso is a new library for sparse learning problems in R and Python.
Neural network training is computationally and memory intensive. Sparse training can reduce the burden on emerging hardware platforms designed to accelerate sparse computations, but it can affect network convergence. In this work, we propose a novel CNN training algorithm Sparse Weight Activation Training (SWAT). SWAT …
This research improves neural network performance with adaptive activation functions in sparse data settings.
New method uses temperature to control sparse MoE convergence rates.
New RL agent learns sparse rewards efficiently.
Study confirms sparse coding in whole brain using MRI data.
Regularization improves stability and consistency of sparse autoencoders.
Sparse neural networks can match dense models on Lipschitz functions.
Model shows loss curve with two distinct exponents due to sparse activations.
Concerns about interpretability, computational resources, and principled inductive priors have motivated efforts to engineer sparse neural models for NLP tasks. If sparsity is important for NLP, might well-trained neural models naturally become roughly sparse? Using the Taxi-Euclidean norm to measure sparsity, we find …
Batteryless or so called passive wearables are providing new and innovative methods for human activity recognition (HAR), especially in healthcare applications for older people. Passive sensors are low cost, lightweight, unobtrusive and desirably disposable; attractive attributes for healthcare applications in hospital…
Sparse Transformers degrade semantic information first, with early layers encoding more.
Unified framework for clustering with sparse convex combinations.
Binary autoencoder with sparse hidden layer preserves information and zero reconstruction error.
Transformers exhibit sparse activation maps, reducing computational load and improving robustness.
We investigate sparse representations for control in reinforcement learning. While these representations are widely used in computer vision, their prevalence in reinforcement learning is limited to sparse coding where extracting representations for new data can be computationally intensive. Here, we begin by demonstrat…
In this paper, we consider the problem of recovering a sparse signal based on penalized least squares formulations. We develop a novel algorithm of primal-dual active set type for a class of nonconvex sparsity-promoting penalties, including , bridge, smoothly clipped absolute deviation, capped and mini…
SAEs struggle with curved activation manifolds, revealing layer-dependent scaling laws.
A method to improve sequential learning by keeping past data errors in check.
We propose a sparse-coding framework for activity recognition in ubiquitous and mobile computing that alleviates two fundamental problems of current supervised learning approaches. (i) It automatically derives a compact, sparse and meaningful feature representation of sensor data that does not rely on prior expert know…
Deep Convolutional Sparse Coding (D-CSC) is a framework reminiscent of deep convolutional neural networks (DCNNs), but by omitting the learning of the dictionaries one can more transparently analyse the role of the activation function and its ability to recover activation paths through the layers. Papyan, Romano, and E…
Sparse Canonical Correlation Analysis (CCA) has received considerable attention in high-dimensional data analysis to study the relationship between two sets of random variables. However, there has been remarkably little theoretical statistical foundation on sparse CCA in high-dimensional settings despite active methodo…
DiAL uses Bayesian Dirichlet random fields for active learning with sparse labels.
New framework tackles high-dimensional reliability analysis using surrogate models and active subspaces.
The paper provides statistical guarantees for sparse deep learning.
The paper improves Gaussian process models for efficient batch optimization.
Sparse Subspace Clustering (SSC) is a state-of-the-art method for clustering high-dimensional data points lying in a union of low-dimensional subspaces. However, while optimization-based SSC algorithms suffer from high computational complexity, other variants of SSC, such as Orthogonal Matching Pursuit-based S…
Active learning selects optimal measurement times for inferring continuous paths from sparse data.
We propose to execute deep neural networks (DNNs) with dynamic and sparse graph (DSG) structure for compressive memory and accelerative execution during both training and inference. The great success of DNNs motivates the pursuing of lightweight models for the deployment onto embedded devices. However, most of the prev…
We propose a novel general algorithm LHAC that efficiently uses second-order information to train a class of large-scale l1-regularized problems. Our method executes cheap iterations while achieving fast local convergence rate by exploiting the special structure of a low-rank matrix, constructed via quasi-Newton approx…
We study the problem of efficient PAC active learning of homogeneous linear classifiers (halfspaces) in , where the goal is to learn a halfspace with low error using as few label queries as possible. Under the extra assumption that there is a -sparse halfspace that performs well on the data ()…
The digital revolution of the banking system with evolving European regulations have pushed the major banking actors to innovate by a newly use of their clients' digital information. Given highly sparse client activities, we propose CPOPT-Net, an algorithm that combines the CP canonical tensor decomposition, a multidim…
Brain networks in fMRI are typically identified using spatial independent component analysis (ICA), yet mathematical constraints such as sparse coding and positivity both provide alternate biologically-plausible frameworks for generating brain networks. Non-negative Matrix Factorization (NMF) would suppress negative BO…
We study sparse group Lasso for high-dimensional double sparse linear regression, where the parameter of interest is simultaneously element-wise and group-wise sparse. This problem is an important instance of the simultaneously structured model -- an actively studied topic in statistics and machine learning. In the noi…
A new method for the unsupervised learning of sparse representations using autoencoders is proposed and implemented by ordering the output of the hidden units by their activation value and progressively reconstructing the input in this order. This can be done efficiently in parallel with the use of cumulative sums and …
Finding sparse solutions of underdetermined systems of linear equations is a fundamental problem in signal processing and statistics which has become a subject of interest in recent years. In general, these systems have infinitely many solutions. However, it may be shown that sufficiently sparse solutions may be identi…
New CSC model extracts EEG signals with low noise sensitivity.
The paper analyzes dynamics of momentum in high dimensions with sparse updates.
New algorithms allow multiple robots to search efficiently without central coordination.
o1Neuro neural network approximates complex functions and converges quickly.
Most artificial networks today rely on dense representations, whereas biological networks rely on sparse representations. In this paper we show how sparse representations can be more robust to noise and interference, as long as the underlying dimensionality is sufficiently high. A key intuition that we develop is that …