Dropout speeds up convergence in shallow linear NNs, with a rate bound.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper characterizes functions of shallow ReLU NN denoisers under minimal norm constraints.
This paper explores how kernel methods can explain data effects on neural collapse.
We reparametrize ReLU NNs as splines to understand their learning dynamics.
We propose a one-class neural network (OC-NN) model to detect anomalies in complex data sets. OC-NN combines the ability of deep networks to extract a progressively rich representation of data with the one-class objective of creating a tight envelope around normal data. The OC-NN approach breaks new ground for the foll…
Recent theoretical work has demonstrated that deep neural networks have superior performance over shallow networks, but their training is more difficult, e.g., they suffer from the vanishing gradient problem. This problem can be typically resolved by the rectified linear unit (ReLU) activation. However, here we show th…
The paper certifies neural network-based control barrier functions efficiently.
Paper studies the expressivity of Convolutional Neural Networks (CNNs).
Wide neural networks can benefit from multi-task learning in their infinite-width limit.
Shallow water environments create a challenging channel for communications. In this paper, we focus on the challenges posed by the frequency-selective signal distortion called the Doppler effect. We explore the design and performance of machine learning (ML) based demodulation methods --- (1) Deep Belief Network-feed f…
Neural networks estimate statistical divergences with performance guarantees.
Wide neural networks converge linearly to zero loss with feature learning.
We present a neural network (NN) approach to fit and predict implied volatility surfaces (IVSs). Atypically to standard NN applications, financial industry practitioners use such models equally to replicate market prices and to value other financial instruments. In other words, low training losses are as important as g…
This paper characterizes how randomized neural networks generalize well in multi-dimensional tasks.
In this work, we propose a new training method for finding minimum weight norm solutions in over-parameterized neural networks (NNs). This method seeks to improve training speed and generalization performance by framing NN training as a constrained optimization problem wherein the sum of the norm of the weights in each…
NeuraLUT maps neural networks to lookup tables, reducing latency and improving expressivity.
This paper extends ResNet theory to infinitely deep networks, linking them to diffusion processes.
To help understand the underlying mechanisms of neural networks (NNs), several groups have, in recent years, studied the number of linear regions of piecewise linear functions generated by deep neural networks (DNN). In particular, they showed that can grow exponentially with the number of network paramet…
Neural networks learn higher-order derivatives for physics problems.
Two-layer NN with channel attention learns low-degree spherical polynomials efficiently.
We investigated the feature map inside deep neural networks (DNNs) by tracking the transport map. We are interested in the role of depth (why do DNNs perform better than shallow models?) and the interpretation of DNNs (what do intermediate layers do?) Despite the rapid development in their application, DNNs remain anal…
We propose a simple approach which, given distributed computing resources, can nearly achieve the accuracy of -NN prediction, while matching (or improving) the faster prediction time of -NN. The approach consists of aggregating denoised -NN predictors over a small number of distributed subsamples. We show, bot…
Randomly trained neural networks can generalize well if there's a simpler underlying teacher model.
Study bridges GARCH and NN models for volatility forecasting.
This paper analyzes deep Stable neural networks, showing convergence rates under different growth settings.
We analyze symmetries in overparametrized neural networks using a mean-field approach.
A fast method for LOOCV in k-NN regression reduces computation time.
Adds layers to NNs to protect them from reverse engineering.
Artificial neural networks (NN) are instrumental in realizing highly-automated driving functionality. An overarching challenge is to identify best safety engineering practices for NN and other learning-enabled components. In particular, there is an urgent need for an adequate set of metrics for measuring all-important …
Prototype rules simplify multiclass classification in metric spaces, achieving consistency and reduced complexity.
Neural Networks (NN) have recently emerged as backbone of several sensitive applications like automobile, medical image, security, etc. NNs inherently offer Partial Fault Tolerance (PFT) in their architecture; however, the biased PFT of NNs can lead to severe consequences in applications like cryptography and security …
Shallow neural networks can represent polynomials efficiently.
This work establishes the equivalence between neural networks and support vector machines.
In the -nearest neighborhood model (-NN), we are given a set of points , and we shall answer queries by returning the nearest neighbors of in according to some metric. This concept is crucial in many areas of data analysis and data processing, e.g., computer vision, document retrieval and machi…
We focus on estimating \emph{a priori} generalization error of two-layer ReLU neural networks (NNs) trained by mean squared error, which only depends on initial parameters and the target function, through the following research line. We first estimate \emph{a priori} generalization error of finite-width two-layer ReLU …
Machine Learning (ML) is making a strong resurgence in tune with the massive generation of unstructured data which in turn requires massive computational resources. Due to the inherently compute- and power-intensive structure of Neural Networks (NNs), hardware accelerators emerge as a promising solution. However, with …
This paper studies large-width asymptotics for ReLU neural networks with α-Stable initializations.
Qualitative analysis of MC dropout for NN model uncertainty.
Training a Neural Network (NN) with lots of parameters or intricate architectures creates undesired phenomena that complicate the optimization process. To address this issue we propose a first modular approach to NN design, wherein the NN is decomposed into a control module and several functional modules, implementing …
The paper optimizes k-NN for distributed learning with minimax optimal performance.
Despite the success of neural networks (NNs), there is still a concern among many over their "black box" nature. Why do they work? Here we present a simple analytic argument that NNs are in fact essentially polynomial regression models. This view will have various implications for NNs, e.g. providing an explanation for…
Neural Networks (NNs) have been extensively used for a wide spectrum of real-world regression tasks, where the goal is to predict a numerical outcome such as revenue, effectiveness, or a quantitative result. In many such tasks, the point prediction is not enough: the uncertainty (i.e. risk or confidence) of that predic…
Study on dropout in neural networks using percolation theory.
-nearest neighbour (-NN) is one of the simplest and most widely-used methods for supervised classification, that predicts a query's label by taking weighted ratio of observed labels of objects nearest to the query. The weights and the parameter regulate its bias-variance trade-off, and the …
We derive high-probability finite-sample uniform rates of consistency for -NN regression that are optimal up to logarithmic factors under mild assumptions. We moreover show that -NN regression adapts to an unknown lower intrinsic dimension automatically. We then apply the -NN regression rates to establish new …
Wide and shallow networks approximate convex functions well.
Ensembles of neural networks (NNs) have long been used to estimate predictive uncertainty; a small number of NNs are trained from different initialisations and sometimes on differing versions of the dataset. The variance of the ensemble's predictions is interpreted as its epistemic uncertainty. The appeal of ensembling…
A garland based on a manifold is a finite set of manifolds homeomorphic to with some of them glued together at marked points. Fix a manifold and consider a space $\NN$ of all smooth mappings of garlands based on into . We construct operations and on the bordism groups $\bor_*(\NN)$ …