This work shows how penalising bias terms in norm regularisation leads to sparse solutions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper advances sparse regularisation theory for measures with new kernel insights.
Study on GD and SGD over diagonal networks, focusing on stepsizes and regularisation.
A new method for manifold learning using sparse regularised optimal transport.
QSD enhances deep network performance through biologically plausible dropout.
SPReD uses uncertainty to decide imitation from demonstrations.
Learning sparse linear models with two-way interactions is desirable in many application domains such as genomics. l1-regularised linear models are popular to estimate sparse models, yet standard implementations fail to address specifically the quadratic explosion of candidate two-way interactions in high dimensions, a…
In this short report, we discuss how coordinate-wise descent algorithms can be used to solve minimum variance portfolio (MVP) problems in which the portfolio weights are constrained by norms, where . A portfolio which weights are regularised by such norms is called a sparse portfolio (Brodie et …
Recently, metric learning and similarity learning have attracted a large amount of interest. Many models and optimisation algorithms have been proposed. However, there is relatively little work on the generalization analysis of such methods. In this paper, we derive novel generalization bounds of metric and similarity …
We introduce a framework for Continual Learning (CL) based on Bayesian inference over the function space rather than the parameters of a deep neural network. This method, referred to as functional regularisation for Continual Learning, avoids forgetting a previous task by constructing and memorising an approximate post…
The kernel null-space technique and its regression-based formulation (called one-class kernel spectral regression, a.k.a. OC-KSR) is known to be an effective and computationally attractive one-class classification framework. Despite its outstanding performance, the applicability of kernel null-space method is limited d…
The use of L1 regularisation for sparse learning has generated immense research interest, with successful application in such diverse areas as signal acquisition, image coding, genomics and collaborative filtering. While existing work highlights the many advantages of L1 methods, in this paper we find that L1 regularis…
Novel deep learning method predicts reaction coordinates and future MD trajectories.
This paper improves inverse problem solving with weakly convex regularisers and proves convergence.
Deep learning using multi-layer neural networks (NNs) architecture manifests superb power in modern machine learning systems. The trained Deep Neural Networks (DNNs) are typically large. The question we would like to address is whether it is possible to simplify the NN during training process to achieve a reasonable pe…
GNIs induce a regulariser that penalizes high-frequency components in neural network activations.
Single-cell RNA sequencing (scRNA-seq) is a fast growing approach to measure the genome-wide transcriptome of many individual cells in parallel, but results in noisy data with many dropout events. Existing methods to learn molecular signatures from bulk transcriptomic data may therefore not be adapted to scRNA-seq data…
A Python package solves source duplication in single channel LVMs using spectral regularisation.
We consider a general regularised interpolation problem for learning a parameter vector from data. The well known representer theorem says that under certain conditions on the regulariser there exists a solution in the linear span of the data points. This is the core of kernel methods in machine learning as it makes th…
We introduce a class of regularisable infinite dimensional principal fibre bundles which includes fibre bundles arising in gauge field theories like Yang-Mills and string theory and which generalise finite dimensional Riemannian principal fibre bundles induced by an isometric action. We show that the orbits of regulari…
New framework monitors neural network training and reveals regularisation mechanisms.
This work uncovers algorithm-dependent regularisation in diffusion models.
Despite recent advances in regularisation theory, the issue of parameter selection still remains a challenge for most applications. In a recent work the framework of statistical learning was used to approximate the optimal Tikhonov regularisation parameter from noisy data. In this work, we improve their results and ext…
One of the fundamental tasks of science is to find explainable relationships between observed phenomena. One approach to this task that has received attention in recent years is based on probabilistic graphical modelling with sparsity constraints on model structures. In this paper, we describe two new approaches to Bay…
The problem of adversarial examples has highlighted the need for a theory of regularisation that is general enough to apply to exotic function classes, such as universal approximators. In response, we give a very general equality result regarding the relationship between distributional robustness and regularisation, as…
Meta-learning improves support recovery in high-dimensional PCA.
New method prevents deep learning forgetting past by remembering key examples.
Adaptive, sparse graphs improve learning performance.
Bayesian framework for encoding uncertainty and inducing sparsity.
Novel approach finds implicit regularisation in two-player games using BEA.
Regularization preserves topological data structure in autoencoders.
Regularised canonical correlation analysis was recently extended to more than two sets of variables by the multiblock method Regularised generalised canonical correlation analysis (RGCCA). Further, Sparse GCCA (SGCCA) was proposed to address the issue of variable selection. However, for technical reasons, the variable …
Investigates gradient descent dynamics and introduces new regularisation methods.
Researchers compare different gradient methods for ridge regression, finding conjugate gradients have similar performance.
SPARTAN learns sparse interaction graphs between objects in scenes.
We propose graph-dependent implicit regularisation strategies for distributed stochastic subgradient descent (Distributed SGD) for convex problems in multi-agent learning. Under the standard assumptions of convexity, Lipschitz continuity, and smoothness, we establish statistical learning rates that retain, up to logari…
Regularizes ML algorithms for robust multivariate analysis against distribution shifts.
Study on how noise and variation-norm regularisation help shallow ReLU networks use fewer neurons.
New algorithm STCV improves sparse model discovery from normalised data.
This work optimizes RL algorithms using entropy regularisation for continuous-time LQ problems.
Paper studies particle method for LSV model calibration, proving convergence and error bounds.
This study uses continuous-time analysis to understand how momentum affects the optimisation of diagonal linear networks.
Logistic regression models with observations and linearly-independent covariates are shown to have Fisher information volumes which are bounded below by and above by . This is proved with a novel generalization of the classical theorems of Pythagoras and de Gua, which is of independent …
Effective regularisation of neural networks is essential to combat overfitting due to the large number of parameters involved. We present an empirical analogue to the Lipschitz constant of a feed-forward neural network, which we refer to as the maximum gain. We hypothesise that constraining the gain of a network will h…
New insights into how neural networks learn features, especially when they are very wide.
Study on neural networks with regularisation and its impact on training dynamics.
Recent work has established the equivalence between deep neural networks and Gaussian processes (GPs), resulting in so-called neural network Gaussian processes (NNGPs). The behaviour of these models depends on the initialisation of the corresponding network. In this work, we consider the impact of noise regularisation …
Matrix factorisation methods decompose multivariate observations as linear combinations of latent feature vectors. The Indian Buffet Process (IBP) provides a way to model the number of latent features required for a good approximation in terms of regularised reconstruction error. Previous work has focussed on latent fe…