We find a convex model for traditional nonlinear regression under L2 loss.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Adversarial training makes logistic regression weight loss landscapes sharper.
Dynamic-weight AMMs outperform traditional CEX rebalancing in tokenized funds, especially on L2s.
We give a topological interpretation of the space of L2-harmonic forms on finite-volume manifolds with sufficiently pinched negative curvature. We give examples showing that this interpretation fails if the curvature is not sufficiently pinched and that our result is sharp with respect to the pinching constants. The me…
Paper proves L2 regression can learn k-juntas without distributional assumptions.
Batch Normalization is a commonly used trick to improve the training of deep neural networks. These neural networks use L2 regularization, also called weight decay, ostensibly to prevent overfitting. However, we show that L2 regularization has no regularizing effect when combined with normalization. Instead, regulariza…
The paper connects neural collapse and low-rank bias in networks with L2 regularization.
Improves deep learning performance on noisy datasets using inverse-variance weighting.
Regularization improves stability and consistency of sparse autoencoders.
Transfer learning through fine-tuning a pre-trained neural network with an extremely large dataset, such as ImageNet, can significantly accelerate training while the accuracy is frequently bottlenecked by the limited dataset size of the new target task. To solve the problem, some regularization methods, constraining th…
We present a generalization of the Cauchy/Lorentzian, Geman-McClure, Welsch/Leclerc, generalized Charbonnier, Charbonnier/pseudo-Huber/L1-L2, and L2 loss functions. By introducing robustness as a continuous parameter, our loss function allows algorithms built around robust loss minimization to be generalized, which imp…
We study a family of sparse estimators defined as minimizers of some empirical Lipschitz loss function -- which include the hinge loss, the logistic loss and the quantile regression loss -- with a convex, sparse or group-sparse regularization. In particular, we consider the L1 norm on the coefficients, its sorted Slope…
The main theme of this work is a unifying algorithm, \textbf{L}oop\textbf{L}ess \textbf{S}ARAH (L2S) for problems formulated as summation of individual loss functions. L2S broadens a recently developed variance reduction method known as SARAH. To find an -accurate solution, L2S enjoys a complexity of ${\cal O}\b…
Paper introduces stability in model averaging and proposes a L2-penalty method.
Develops a new fuzzy model using QPs and ewl2 regularization to improve local region behavior.
We give a fast oblivious L2-embedding of to satisfying Our embedding dimension equals , a constant independent of the distortion . We use as a black-box any L2-embedding $Π…
In this paper, we consider one dimensional (shallow) ReLU neural networks in which weights are chosen randomly and only the terminal layer is trained. First, we mathematically show that for such networks L2-regularized regression corresponds in function space to regularizing the estimate's second derivative for fairly …
Improved sample efficiency in learning sparse Ising models.
New properties of weighted Hilbert transform derived, useful for imaging applications.
The use of machine-learning in neuroimaging offers new perspectives in early diagnosis and prognosis of brain diseases. Although such multivariate methods can capture complex relationships in the data, traditional approaches provide irregular (l2 penalty) or scattered (l1 penalty) predictive pattern with a very limited…
Importance-weighted risk minimization is a key ingredient in many machine learning algorithms for causal inference, domain adaptation, class imbalance, and off-policy reinforcement learning. While the effect of importance weighting is well-characterized for low-capacity misspecified models, little is known about how it…
Transfer learning have been frequently used to improve deep neural network training through incorporating weights of pre-trained networks as the starting-point of optimization for regularization. While deep transfer learning can usually boost the performance with better accuracy and faster convergence, transferring wei…
Distributional (or distribution-valued) data are a new type of data arising from several sources and are considered as realizations of distributional variables. A new set of fuzzy c-means algorithms for data described by distributional variables is proposed. The algorithms use the Wasserstein distance between dist…
ESNs trained with Tikhonov least squares approximate ergodic dynamical systems in L2(μ) norm.
Weight normalization and reparametrized gradient descent adaptively regularize weights and converge to minimum l2 norm solutions.
Proposes TCV for selecting models in regions of interest.
Self-adaptive PINNs improve accuracy in stiff PDEs.
In this paper, we study Lichnerowicz type estimate for eigenvalues of drifting Laplacian operator and L1 and L2 energy for drifting heat equation on closed manifolds with weighted measure. In some sense, this study is about the eigenvalue estimate on Ricci solitons.
Conjugate gradient (CG) methods are a class of important methods for solving linear equations and nonlinear optimization problems. In this paper, we propose a new stochastic CG algorithm with variance reduction and we prove its linear convergence with the Fletcher and Reeves method for strongly convex and smooth functi…
COS method convergence conditions expanded for heavy-tailed distributions.
Forecast dam inflow using sea surface feature weights.
Traffic forecasting is a particularly challenging application of spatiotemporal forecasting, due to the time-varying traffic patterns and the complicated spatial dependencies on road networks. To address this challenge, we learn the traffic network as a graph and propose a novel deep learning framework, Traffic Graph C…
Optimizes CNNs by directing gradients along output channels.
In this paper we study deep learning-based music source separation, and explore using an alternative loss to the standard spectrogram pixel-level L2 loss for model training. Our main contribution is in demonstrating that adding a high-level feature loss term, extracted from the spectrograms using a VGG net, can improve…
Optimizes pruning masks for neural networks using probabilistic fine-tuning and PAC-Bayes bounds.
A standing conjecture in L2-cohomology is that every finite CW-complex X is of L2-determinant class. In this paper, we prove this whenever the fundamental group belongs to a large class of groups containing e.g. all extensions of residually finite groups with amenable quotients, all residually amenable groups and free …
Feature normalization prevents collapse in non-contrastive learning dynamics.
In this paper, we propose a novel linear discriminant analysis criterion via the Bhattacharyya error bound estimation based on a novel L1-norm (L1BLDA) and L2-norm (L2BLDA). Both L1BLDA and L2BLDA maximize the between-class scatters which are measured by the weighted pairwise distances of class means and meanwhile mini…
For a normal covering over a closed oriented topological manifold we give a proof of the L2-signature theorem with twisted coefficients, using Lipschitz structures and the Lipschitz signature operator introduced by Teleman. We also prove that the L-theory isomorphism conjecture as well as the C^*_max-version of the Bau…
Study extends holomorphic forms on noncompact Kahler manifolds.
Dropout is one of the key techniques to prevent the learning from overfitting. It is explained that dropout works as a kind of modified L2 regularization. Here, we shed light on the dropout from Bayesian standpoint. Bayesian interpretation enables us to optimize the dropout rate, which is beneficial for learning of wei…
We provide a proof for an inequality between volume and L2-Betti numbers of aspherical manifolds for which Gromov outlined a strategy based on general ideas of Connes. The implementation of that strategy involves measured equivalence relations, Gaboriau's theory of L2-Betti numbers of R-simplicial complexes, and other …
New method improves convergence of spatial filters in neural networks.
Paper quantifies MEV on L2 networks, finding significant amounts on Polygon.
We prove that L2-Boosting lacks a theoretical property which is central to the behaviour of l1-penalized methods such as basis pursuit and the Lasso: Whereas l1-penalized methods are guaranteed to recover the sparse parameter vector in a high-dimensional linear model under an appropriate restricted nullspace property, …
New L2 regularization improves softmax MAB performance.
Extends L2-norm LDA to 2D inputs using Bhattacharyya bound.
We examine the effect of the Group Lasso (gLasso) regularizer in selecting the salient nodes of Deep Neural Network (DNN) hidden layers by applying a DNN-HMM hybrid speech recognizer to TED Talks speech data. We test two types of gLasso regularization, one for outgoing weight vectors and another for incoming weight vec…