Batch normalization improves deep networks by aligning their decision boundaries with data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Normalizing flows attempt to model an arbitrary probability distribution through a set of invertible mappings. These transformations are required to achieve a tractable Jacobian determinant that can be used in high-dimensional scenarios. The first normalizing flow designs used coupling layer mappings built upon affine …
This note is the updated outline of the article "Interpolational properties of planar spiral curves", Fund. and Applied Math., 2001, Vol.7, N.2, 441-463, published in Russian. The main result establishes boundary regions for spiral and piecewise spiral splines, matching given data. The width of such region can serve as…
Nonlinearity is crucial to the performance of a deep (neural) network (DN). To date there has been little progress understanding the menagerie of available nonlinearities, but recently progress has been made on understanding the rôle played by piecewise affine and convex nonlinearities like the ReLU and absolute value …
We study the geometry of deep (neural) networks (DNs) with piecewise affine and convex nonlinearities. The layers of such DNs have been shown to be {\em max-affine spline operators} (MASOs) that partition their input space and apply a region-dependent affine mapping to their input to produce their output. We demonstrat…
We reparametrize ReLU NNs as splines to understand their learning dynamics.
Neural networks can represent complex piecewise functions efficiently.
We use partial actions, as formalized by Exel, to construct various commensurating actions. We use this in the context of groups piecewise preserving a geometric structure, and we interpret the transfixing property of these commensurating actions as the existence of a model for which the group acts preserving the geome…
BN refines local partition geometry in piecewise-affine networks during training.
In exchange for large quantities of data and processing power, deep neural networks have yielded models that provide state of the art predication capabilities in many fields. However, a lack of strong guarantees on their behaviour have raised concerns over their use in safety-critical applications. A first step to unde…
Paper presents ABGD for efficient piecewise linear regression in high dimensions.
The paper develops a new method for estimating non-parametric regression functions with spatio-temporal dependencies.
Extends gradient-based optimization to spline functions.
We build a rigorous bridge between deep networks (DNs) and approximation theory via spline functions and operators. Our key result is that a large class of DNs can be written as a composition of max-affine spline operators (MASOs), which provide a powerful portal through which to view and analyze their inner workings. …
Improves BN graph learning with splines for scalability.
Deep neural networks (DNNs) generate much richer function spaces than shallow networks. Since the function spaces induced by shallow networks have several approximation theoretic drawbacks, this explains, however, not necessarily the success of deep networks. In this article we take another route by comparing the expre…
Method identifies latent variables from high-dimensional data with piecewise affine mixing.
Regularized least-squares approaches have been successfully applied to linear system identification. Recent approaches use quadratic penalty terms on the unknown impulse response defined by stable spline kernels, which control model space complexity by leveraging regularity and bounded-input bounded-output stability. T…
We propose to optimize the activation functions of a deep neural network by adding a corresponding functional regularization to the cost function. We justify the use of a second-order total-variation criterion. This allows us to derive a general representer theorem for deep neural networks that makes a direct connectio…
We consider the generic regularized optimization problem . Efron, Hastie, Johnstone and Tibshirani [Ann. Statist. 32 (2004) 407--499] have shown that for the LASSO--that is, if is squared error loss and is the norm of --the opti…
Paper develops algorithms for PWA systems with polynomial regret.
We study the coarse geometry of the moduli space of dilation tori with two singularities and the dynamical properties of the action of the Teichmuller flow on this moduli space. This leads to a proof that the vertical foliation of a dilation torus is almost always Morse-Smale. As a corollary, we get that the generic pi…
We study trend filtering, a recently proposed tool of Kim et al. [SIAM Rev. 51 (2009) 339-360] for nonparametric regression. The trend filtering estimate is defined as the minimizer of a penalized least squares criterion, in which the penalty term sums the absolute th order discrete derivatives over the input points…
This study compares different types of normalizing flows for generating complex distributions.
Max-affine regression method converges linearly using GD and SGD.
Study growth patterns in random networks using i.i.d. perturbations.
We connect a large class of Generative Deep Networks (GDNs) with spline operators in order to derive their properties, limitations, and new opportunities. By characterizing the latent space partition, dimension and angularity of the generated manifold, we relate the manifold dimension and approximation error to the sam…
Paper proposes a new method for SP with covariates using PADR and ERM.
PAR provides a flexible framework for quantization in optimization problems.
Many problems on signal processing reduce to nonparametric function estimation. We propose a new methodology, piecewise convex fitting (PCF), and give a two-stage adaptive estimate. In the first stage, the number and location of the change points is estimated using strong smoothing. In the second stage, a constrained s…
ReLU neural networks define piecewise linear functions of their inputs. However, initializing and training a neural network is very different from fitting a linear spline. In this paper, we expand empirically upon previous theoretical work to demonstrate features of trained neural networks. Standard network initializat…
Study reconstructs Faber-Schauder coefficients from antiderivative observations.
This paper considers affine analogues of the isoperimetric inequality in the sense of piecewise linear topology. Given a closed polygon P embedded in R^d having n edges, we give upper and lower bounds for the minimal number of triangles needed to forma triangulated embedded orientable surface in R^d having P as its geo…
New method characterizes surface quadrilateral layouts as special immersions.
The classical approach to linear system identification is given by parametric Prediction Error Methods (PEM). In this context, model complexity is often unknown so that a model order selection step is needed to suitably trade-off bias and variance. Recently, a different approach to linear system identification has been…
Exact LAD line fitting via PALB with linear scaling and speed.
We study additive models built with trend filtering, i.e., additive models whose components are each regularized by the (discrete) total variation of their th (discrete) derivative, for a chosen integer . This results in th degree piecewise polynomial components, (e.g., gives piecewise constant co…
Let G be a connected compact Lie group acting on a manifold M and let D be a transversally elliptic operator on M. The multiplicity of the index of D is a function on the set of irreducible representations of G. Let T be a maximal torus of G with Lie algebra Lie(T). We construct a finite number of piecewise polynomial …
The paper designs neural networks with assurance for controlling nonlinear systems.
PARC uses piecewise linear predictors for regression and classification.
Piecewise polynomial interpolation-based gradient descent reduces oracle complexity for smooth loss functions.
New method uses DC functions for piecewise linear regression.
Kronecker trend filtering improves lattice data smoothing.
We consider the quasiconformal dilatation of projective transformations of the real projective plane. For non-affine transformations, the contour lines of dilatation form a hyperbolic pencil of circles, and these are the only circles that are mapped to circles. We apply this result to analyze the dilatation of the circ…
The paper provides results regarding the computational complexity of hybrid system identification. More precisely, we focus on the estimation of piecewise affine (PWA) maps from input-output data and analyze the complexity of computing a global minimizer of the error. Previous work showed that a global solution could b…
New EM algorithm improves deep generative network training.
Paper finds maximum curvature of Bézier-spline curves.
Efficiently finds sparse solutions to max-plus equations for convex regression.