We reparametrize ReLU NNs as splines to understand their learning dynamics.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Improves BN graph learning with splines for scalability.
Regularized least-squares approaches have been successfully applied to linear system identification. Recent approaches use quadratic penalty terms on the unknown impulse response defined by stable spline kernels, which control model space complexity by leveraging regularity and bounded-input bounded-output stability. T…
We consider the generic regularized optimization problem . Efron, Hastie, Johnstone and Tibshirani [Ann. Statist. 32 (2004) 407--499] have shown that for the LASSO--that is, if is squared error loss and is the norm of --the opti…
This note is the updated outline of the article "Interpolational properties of planar spiral curves", Fund. and Applied Math., 2001, Vol.7, N.2, 441-463, published in Russian. The main result establishes boundary regions for spiral and piecewise spiral splines, matching given data. The width of such region can serve as…
Deep neural networks (DNNs) generate much richer function spaces than shallow networks. Since the function spaces induced by shallow networks have several approximation theoretic drawbacks, this explains, however, not necessarily the success of deep networks. In this article we take another route by comparing the expre…
We propose to optimize the activation functions of a deep neural network by adding a corresponding functional regularization to the cost function. We justify the use of a second-order total-variation criterion. This allows us to derive a general representer theorem for deep neural networks that makes a direct connectio…
Batch normalization improves deep networks by aligning their decision boundaries with data.
ReLU neural networks define piecewise linear functions of their inputs. However, initializing and training a neural network is very different from fitting a linear spline. In this paper, we expand empirically upon previous theoretical work to demonstrate features of trained neural networks. Standard network initializat…
The classical approach to linear system identification is given by parametric Prediction Error Methods (PEM). In this context, model complexity is often unknown so that a model order selection step is needed to suitably trade-off bias and variance. Recently, a different approach to linear system identification has been…
Normalizing flows attempt to model an arbitrary probability distribution through a set of invertible mappings. These transformations are required to achieve a tractable Jacobian determinant that can be used in high-dimensional scenarios. The first normalizing flow designs used coupling layer mappings built upon affine …
New method characterizes surface quadrilateral layouts as special immersions.
The paper develops a new method for estimating non-parametric regression functions with spatio-temporal dependencies.
We study additive models built with trend filtering, i.e., additive models whose components are each regularized by the (discrete) total variation of their th (discrete) derivative, for a chosen integer . This results in th degree piecewise polynomial components, (e.g., gives piecewise constant co…
Extends gradient-based optimization to spline functions.
Kronecker trend filtering improves lattice data smoothing.
We study trend filtering, a recently proposed tool of Kim et al. [SIAM Rev. 51 (2009) 339-360] for nonparametric regression. The trend filtering estimate is defined as the minimizer of a penalized least squares criterion, in which the penalty term sums the absolute th order discrete derivatives over the input points…
We introduce a variational framework to learn the activation functions of deep neural networks. Our aim is to increase the capacity of the network while controlling an upper-bound of the actual Lipschitz constant of the input-output relation. To that end, we first establish a global bound for the Lipschitz constant of …
Sig-Splines model uses signatures and splines for time series data, achieving universality and convexity.
Cubic spline interpolation on Euclidean space is a standard topic in numerical analysis, with countless applications in science and technology. In several emerging fields, for example computer vision and quantum control, there is a growing need for spline interpolation on curved, non-Euclidean space. The generalization…
Multivariate splines linked to infinitely-wide neural networks with improved numerical performance.
Nonlinearity is crucial to the performance of a deep (neural) network (DN). To date there has been little progress understanding the menagerie of available nonlinearities, but recently progress has been made on understanding the rôle played by piecewise affine and convex nonlinearities like the ReLU and absolute value …
Many problems on signal processing reduce to nonparametric function estimation. We propose a new methodology, piecewise convex fitting (PCF), and give a two-stage adaptive estimate. In the first stage, the number and location of the change points is estimated using strong smoothing. In the second stage, a constrained s…
Study reconstructs Faber-Schauder coefficients from antiderivative observations.
RUMBoost combines RUMs and deep learning for better choice modelling.
The center of a quotient group of piecewise linear homeomorphisms is trivial.
This article provides an attempt to extend concepts from the theory of Riemannian manifolds to piecewise linear spaces. In particular we propose an analogue of the Ricci tensor, which we give the name of an Einstein vector field. On a given set of piecewise linear spaces we define and discuss (normalized) Ricci flows. …
A new spline method for manifold learning using Hessian-based curvature penalties.
Global approximation for piecewise linear paths via signatures.
This paper develops a new method for constructing splines on Lie groups using Poisson equation solutions.
PARC uses piecewise linear predictors for regression and classification.
Investigates stability of piecewise flat Ricci flow using analysis and simulations.
We prove that every piecewise linear manifold of dimension up to four on which a finite group acts by piecewise linear homeomorphisms admits a compatible smooth structure with respect to which the group acts smoothly. This solves a challenge posed by Thurston in dimension three and confirms a conjecture by Kwasik and L…
Let G be a connected compact Lie group acting on a manifold M and let D be a transversally elliptic operator on M. The multiplicity of the index of D is a function on the set of irreducible representations of G. Let T be a maximal torus of G with Lie algebra Lie(T). We construct a finite number of piecewise polynomial …
New framework explains deep neural networks using variational spline theory.
Existing dimensionality reduction methods are adept at revealing hidden underlying manifolds arising from high-dimensional data and thereby producing a low-dimensional representation. However, the smoothness of the manifolds produced by classic techniques over sparse and noisy data is not guaranteed. In fact, the embed…
Temporal Functional Circuits explain KAN forecasts with interpretable edge functions.
Piecewise linear activations create many spurious local minima in neural networks.
Piecewise polynomial interpolation-based gradient descent reduces oracle complexity for smooth loss functions.
Signature uniquely identifies piecewise linear surfaces up to thin homotopy.
We study a novel spline-like basis, which we name the "falling factorial basis", bearing many similarities to the classic truncated power basis. The advantage of the falling factorial basis is that it enables rapid, linear-time computations in basis matrix multiplication and basis matrix inversion. The falling factoria…
Dropout improves regularization in flexible models for rare features.
The paper extends a variance gamma model to quadratic functions, reducing arbitrage and computational costs.
In this paper, we introduce a bordism category whose objects are bundles of closed -dimensional piecewise linear manifolds and whose morphisms are bundles of -dimensional piecewise linear cobordisms. In the main theorem of this article, we show that the classifying space $B\mathcal{C}_d^{…
First explicit isometric immersion of a flat Klein bottle in 3D space.
GraN-GAN normalizes gradients for better GAN performance.
Paper proposes variational inference for piecewise-linear systems.
A normalizing flow models a complex probability density as an invertible transformation of a simple density. The invertibility means that we can evaluate densities and generate samples from a flow. In practice, autoregressive flow-based models are slow to invert, making either density estimation or sample generation sl…