Deep neural networks with piecewise-polynomial activations can approximate smooth functions and their derivatives.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper analyzes a simple neural network model with algebraic methods.
Study extends GNN VC dimension bounds to Pfaffian activation functions.
This work interprets GELU and related activations via a first-order loss function.
In this note we consider setups in which variational objectives for Bayesian neural networks can be computed in closed form. In particular we focus on single-layer networks in which the activation function is piecewise polynomial (e.g. ReLU). In this case we show that for a Normal likelihood and structured Normal varia…
The paper approximates Levi-Civita connection and curvature on 2D manifolds using finite elements.
New algorithm reduces dynamic regret for noisy gradient feedback with piecewise polynomial comparators.
SURF simplifies distribution estimation with simple, robust, and fast algorithms.
Piecewise polynomial interpolation-based gradient descent reduces oracle complexity for smooth loss functions.
This paper introduces a novel mixture model-based approach for simultaneous clustering and optimal segmentation of functional data which are curves presenting regime changes. The proposed model consists in a finite mixture of piecewise polynomial regression models. Each piecewise polynomial regression model is associat…
PolyLUT uses polynomials to reduce FPGA latency.
Proposed by Donoho (1997), Dyadic CART is a nonparametric regression method which computes a globally optimal dyadic decision tree and fits piecewise constant functions in two dimensions. In this article we define and study Dyadic CART and a closely related estimator, namely Optimal Regression Tree (ORT), in the contex…
We give a highly efficient "semi-agnostic" algorithm for learning univariate probability distributions that are well approximated by piecewise polynomial density functions. Let be an arbitrary distribution over an interval which is -close (in total variation distance) to an unknown probability distribution $…
Neural networks with ReLU^k approximate Sobolev functions efficiently via Radon transform.
Constructs finite element spaces for -forms, excluding one subspace.
The paper approximates Einstein tensor using finite elements.
This note is an addendum to our earlier work \cite{humi}. In \cite{humi}, we studied a Hamiltonian action for a generalized Calabi-Yau manifold and showed that the Duistermaat-Heckman theorem holds. The purpose of this note is to show that the density function of the Duistermaa-Heckman measure is a piecewise polynomial…
New algorithm predicts piecewise regular functions online.
Let G be a connected compact Lie group acting on a manifold M and let D be a transversally elliptic operator on M. The multiplicity of the index of D is a function on the set of irreducible representations of G. Let T be a maximal torus of G with Lie algebra Lie(T). We construct a finite number of piecewise polynomial …
New method for robust learning from batches, even adversarial ones.
For any positive integer , there exist neural networks with layers, nodes per layer, and distinct parameters which can not be approximated by networks with layers unless they are exponentially large --- they must possess nodes. This result is proved here for a class o…
Paper proposes algorithms to accurately identify breakpoints in piecewise regression.
Normalizing flows attempt to model an arbitrary probability distribution through a set of invertible mappings. These transformations are required to achieve a tractable Jacobian determinant that can be used in high-dimensional scenarios. The first normalizing flow designs used coupling layer mappings built upon affine …
We consider the fundamental learning problem of estimating properties of distributions over large domains. Using a novel piecewise-polynomial approximation technique, we derive the first unified methodology for constructing sample- and time-efficient estimators for all sufficiently smooth, symmetric and non-symmetric, …
We consider Bayesian analysis of a class of multiple changepoint models. While there are a variety of efficient ways to analyse these models if the parameters associated with each segment are independent, there are few general approaches for models where the parameters are dependent. Under the assumption that the depen…
Extends gradient-based optimization to spline functions.
While all kinds of mixed data -from personal data, over panel and scientific data, to public and commercial data- are collected and stored, building probabilistic graphical models for these hybrid domains becomes more difficult. Users spend significant amounts of time in identifying the parametric form of the random va…
POUnets combine partitions of unity and monomials for efficient deep learning.
Method estimates observation functions in state-space models without supervision.
Finite element method approximates scalar curvature in arbitrary dimensions.
The paper calculates volumes of moduli spaces of flat metrics on spheres with specific angles.
General lower bounds on neural network approximation in L^p norm.
We provide a differentially private algorithm for hypothesis selection. Given samples from an unknown probability distribution and a set of probability distributions , the goal is to output, in a -differentially private manner, a distribution from whose total variation di…
We propose to use deep neural networks for generating samples in Monte Carlo integration. Our work is based on non-linear independent components estimation (NICE), which we extend in numerous ways to improve performance and enable its application to integration problems. First, we introduce piecewise-polynomial couplin…
Many functions of interest are in a high-dimensional space but exhibit low-dimensional structures. This paper studies regression of a -Hölder function in which varies along a central subspace of dimension while . A direct approximation of in with an acc…
We study additive models built with trend filtering, i.e., additive models whose components are each regularized by the (discrete) total variation of their th (discrete) derivative, for a chosen integer . This results in th degree piecewise polynomial components, (e.g., gives piecewise constant co…
Unified analysis of kernel-based and locally adaptive bandit optimization methods.
Many activation functions have been proposed in the past, but selecting an adequate one requires trial and error. We propose a new methodology of designing activation functions within a neural network at each layer. We call this technique an "activation ensemble" because it allows the use of multiple activation functio…
This paper studies activation sparsity in large language models, finding key trends and implications.
Efficiently learns mixtures of Gaussians without separation assumptions.
Activity recognition from sensor data deals with various challenges, such as overlapping activities, activity labeling, and activity detection. Although each challenge in the field of recognition has great importance, the most important one refers to online activity recognition. The present study tries to use online hi…
Many neural network architectures rely on the choice of the activation function for each hidden layer. Given the activation function, the neural network is trained over the bias and the weight parameters. The bias catches the center of the activation, and the weights capture the scale. Here we propose to train the netw…
BinaryDuo improves BNNs by coupling binary activations, outperforming state-of-the-art models.
Evolutionary algorithms improve neural network performance by discovering better activation functions.
Active learning method balances bias and variance under class imbalance.
Probabilistic representations, such as Bayesian and Markov networks, are fundamental to much of statistical machine learning. Thus, learning probabilistic representations directly from data is a deep challenge, the main computational bottleneck being inference that is intractable. Tractable learning is a powerful new p…
Sampling logconcave functions arising in statistics and machine learning has been a subject of intensive study. Recent developments include analyses for Langevin dynamics and Hamiltonian Monte Carlo (HMC). While both approaches have dimension-independent bounds for the underlying processes under s…
The field of statistical relational learning aims at unifying logic and probability to reason and learn from data. Perhaps the most successful paradigm in the field is probabilistic logic programming: the enabling of stochastic primitives in logic programming, which is now increasingly seen to provide a declarative bac…