Online Normalization is a new technique for normalizing the hidden activations of a neural network. Like Batch Normalization, it normalizes the sample dimension. While Online Normalization does not use batches, it is as accurate as Batch Normalization. We resolve a theoretical limitation of Batch Normalization by intro…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper reviews normalization techniques for DNNs.
One of the challenges in the study of generative adversarial networks is the instability of its training. In this paper, we propose a novel weight normalization technique called spectral normalization to stabilize the training of the discriminator. Our new normalization technique is computationally light and easy to in…
Normalization techniques have only recently begun to be exploited in supervised learning tasks. Batch normalization exploits mini-batch statistics to normalize the activations. This was shown to speed up training and result in better models. However its success has been very limited when dealing with recurrent neural n…
A new method normalizes activations to match batch normalization without batch dependence.
We introduce a new normalization technique that exhibits the fast convergence properties of batch normalization using a transformation of layer weights instead of layer outputs. The proposed technique keeps the contribution of positive and negative weights to the layer output balanced. We validate our method on a set o…
This paper explores normalization in neural ODEs, achieving high accuracy in CIFAR-10.
In this work, we investigate Batch Normalization technique and propose its probabilistic interpretation. We propose a probabilistic model and show that Batch Normalization maximazes the lower bound of its marginalized log-likelihood. Then, according to the new probabilistic model, we design an algorithm which acts cons…
Proposes a normalization technique for manifold valued data.
Deep convolutional neural networks are known to be unstable during training at high learning rate unless normalization techniques are employed. Normalizing weights or activations allows the use of higher learning rates, resulting in faster convergence and higher test accuracy. Batch normalization requires minibatch sta…
Training state-of-the-art, deep neural networks is computationally expensive. One way to reduce the training time is to normalize the activities of the neurons. A recently introduced technique called batch normalization uses the distribution of the summed input to a neuron over a mini-batch of training cases to compute…
A virtual link diagram is called normal if the associated abstract link diagram is checkerboard colorable, and a virtual link is normal if it has a normal diagram as a representative.In this paper, we introduce a method of converting a virtual link diagram to a normal virtual link diagram by use of the double covering …
CBN improves batch normalization for small mini-batch sizes.
In this article, we first describe a normal form of real-analytic, Levi-nondegenerate submanifolds of of codimension d 1 under the action of formal biholomorphisms, that is, of perturbations of Levi-nondegenerate hyperquadrics. We give a sufficient condition on the formal normal form that ensures that the n…
We present a new and shorter proof of Stocking's result that any strongly irreducible Heegaard surface of a closed orientable triangulated 3-manifold is isotopic to an almost normal surface. We also re-prove a result of Jaco and Rubinstein on normal spheres. Both proofs are based on the "reduction" technique introduced…
Normalization methods are a central building block in the deep learning toolbox. They accelerate and stabilize training, while decreasing the dependence on manually tuned learning rate schedules. When learning from multi-modal distributions, the effectiveness of batch normalization (BN), arguably the most prominent nor…
While the authors of Batch Normalization (BN) identify and address an important problem involved in training deep networks-- Internal Covariate Shift-- the current solution has certain drawbacks. Specifically, BN depends on batch statistics for layerwise input normalization during training which makes the estimates of …
Deep learning improves GW signal detection efficiency and robustness.
We present a short complete proof of the existence of the normal cycle of a compact subanalytic set. The approach is inspired by some old ides of Joseph Fu, uses Morse theoretic techniques and -minimal topology.
For the last few years it has been observed that the Deep Neural Networks (DNNs) has achieved an excellent success in image classification, speech recognition. But DNNs are suffer great deal of challenges for time series forecasting because most of the time series data are nonlinear in nature and highly dynamic in beha…
Develops an oblique projection technique to approximate a foliation for non-normal dynamics.
In this paper, we study normal homogeneous Finsler spaces. We first define the notion of a normal homogeneous Finsler space, using the method of isometric submersion of Finsler metrics. Then we study the geometric properties. In particular, we establish a technique to reduce the classification of normal homogeneous Fin…
New method builds complex networks from attribute interactions without normalization.
This review compares various deep generative models.
PL-MCMC samples from normalizing flows' conditional distributions.
Spectral normalization stabilizes GANs by controlling gradient explosion and vanishing.
We propose a simple but effective multi-source domain generalization technique based on deep neural networks by incorporating optimized normalization layers that are specific to individual domains. Our approach employs multiple normalization methods while learning separate affine parameters per domain. For each domain,…
New proof of Alesker's Irreducibility Theorem using localization techniques.
A new robust scaling approach improves downstream metabolomics analysis.
Proposes a new normalization method using convolutional neural networks.
Weight normalization speeds up matrix sensing problems.
We consider graphs Sigma^n in R^m with prescribed mean curvature and flat normal bundle. Using techniques of Schoen, Simon and Yau, and Ecker-Huisken, we derive an interior curvature estimate of the form |A|^2<=C/R^2 up to dimension n<=5, where C is a constant depending on natural geometric data of Sigma^n only. This g…
Computational knot theory and 3-manifold topology have seen significant breakthroughs in recent years, despite the fact that many key algorithms have complexity bounds that are exponential or greater. In this setting, experimentation is essential for understanding the limits of practicality, as well as for gauging the …
Calculation of the log-normalizer is a major computational obstacle in applications of log-linear models with large output spaces. The problem of fast normalizer computation has therefore attracted significant attention in the theoretical and applied machine learning literature. In this paper, we analyze a recently pro…
We improve neural network explainability by bypassing batch normalization.
Spectral clustering is a technique that clusters elements using the top few eigenvectors of their (possibly normalized) similarity matrix. The quality of spectral clustering is closely tied to the convergence properties of these principal eigenvectors. This rate of convergence has been shown to be identical for both th…
Batch Normalization (BN) is one of the most widely used techniques in Deep Learning field. But its performance can awfully degrade with insufficient batch size. This weakness limits the usage of BN on many computer vision tasks like detection or segmentation, where batch size is usually small due to the constraint of m…
In this paper we introduce a novel method of gradient normalization and decay with respect to depth. Our method leverages the simple concept of normalizing all gradients in a deep neural network, and then decaying said gradients with respect to their depth in the network. Our proposed normalization and decay techniques…
Improves data normality with robust transformations.
Improved binning technique boosts nUV measure performance.
In this work we investigate the reasons why Batch Normalization (BN) improves the generalization performance of deep networks. We argue that one major reason, distinguishing it from data-independent normalization methods, is randomness of batch statistics. This randomness appears in the parameters rather than in activa…
A new technique normalizes nodes within groups to improve GNN performance.
In recent years, data have become increasingly higher dimensional and, therefore, an increased need has arisen for dimension reduction techniques for clustering. Although such techniques are firmly established in the literature for multivariate data, there is a relative paucity in the area of matrix variate, or three-w…
Traditionally, multi-layer neural networks use dot product between the output vector of previous layer and the incoming weight vector as the input to activation function. The result of dot product is unbounded, thus increases the risk of large variance. Large variance of neuron makes the model sensitive to the change o…
We study the long time behavior of the volume preserving -flow in for . By extending Andrews' technique for the flow along the affine normal, we prove that every centrally symmetric solution to the volume preserving -flow converges sequentially to the unit ball in the $…
EvoMSN tackles time series forecasting under distribution shifts by evolving multi-scale normalization.
Layer normalization improves federated learning with skewed labels.
The aim of this paper is to prove a normal form Theorem for Dirac-Jacobi bundles using the recent techniques from Bursztyn, Lima and Meinrenken. As the most important consequence, we can prove the splitting theorems of Jacobi pairs which was proposed by Dazord, Lichnerowicz and Marle. As an application we provide a alt…