Combines variational autoencoders with normalizing flows for faster training.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The enumeration of normal surfaces is a crucial but very slow operation in algorithmic 3-manifold topology. At the heart of this operation is a polytope vertex enumeration in a high-dimensional space (standard coordinates). Tollefson's Q-theory speeds up this operation by using a much smaller space (quadrilateral coord…
A new algorithm speeds up elliptical slice sampling for truncated multivariate normals.
We present a new similarity measure based on information theoretic measures which is superior than Normalized Compression Distance for clustering problems and inherits the useful properties of conditional Kolmogorov complexity. We show that Normalized Compression Dictionary Size and Normalized Compression Dictionary En…
Preconditioned NFs speed up sampling from complex posterior distributions in inverse problems.
Four new methods for computing generalized chi-square distribution.
We consider embedded hypersurfaces evolving by fully nonlinear flows in which the normal speed of motion is a homogeneous degree one, concave or convex function of the principal curvatures, and prove a non-collapsing estimate: Precisely, the function which gives the curvature of the largest interior sphere touching the…
We study the motion of smooth, strictly convex bodies in expanding in the direction of their normal vector field with speed depending on Gauss curvature and support function.
This paper analyzes how normalization layers improve neural network training.
FedGradNorm improves federated MTL by balancing task learning speeds.
The paper proves convergence of certain curvature flows to the origin.
CNNs improve wind speed forecasts in the Netherlands.
Normalization techniques have only recently begun to be exploited in supervised learning tasks. Batch normalization exploits mini-batch statistics to normalize the activations. This was shown to speed up training and result in better models. However its success has been very limited when dealing with recurrent neural n…
Weight normalization speeds up matrix sensing problems.
New analysis shows surprising results on adaptation speed of causal models.
A new method speeds up spectral normalization for neural nets.
We show that strictly convex surfaces contracting with normal velocity equal to |A|^2 shrink to a point in finite time. After appropriate rescaling, they converge to spheres. We indicate how we used a computer to find the main test function.
Recent seminal work at the intersection of deep neural networks practice and random matrix theory has linked the convergence speed and robustness of these networks with the combination of random weight initialization and nonlinear activation function in use. Building on those principles, we introduce a process to trans…
MTFL improves UA and speeds convergence in personalised DNNs on edge devices.
Convexity preserved in curved surfaces moving at concave speeds.
The normalized maximized likelihood (NML) provides the minimax regret solution in universal data compression, gambling, and prediction, and it plays an essential role in the minimum description length (MDL) method of statistical modeling and estimation. Here we show that the normalized maximum likelihood has a Bayes-li…
Very deep CNNs achieve state-of-the-art results in both computer vision and speech recognition, but are difficult to train. The most popular way to train very deep CNNs is to use shortcut connections (SC) together with batch normalization (BN). Inspired by Self- Normalizing Neural Networks, we propose the self-normaliz…
One of the difficulties of training deep neural networks is caused by improper scaling between layers. Scaling issues introduce exploding / gradient problems, and have typically been addressed by careful scale-preserving initialization. We investigate the value of preserving scale, or isometry, beyond the initial weigh…
In this work, we propose a novel technique to boost training efficiency of a neural network. Our work is based on an excellent idea that whitening the inputs of neural networks can achieve a fast convergence speed. Given the well-known fact that independent components must be whitened, we introduce a novel Independent-…
HNHN learns from hypergraphs with hyperedge neurons for better classification.
Generalizing results of Chou and Wang \cite{1} we study the flows of the leaves of a foliation of consisting of uniformly convex hypersurfaces in the direction of their outer normals with speeds . For quite general functions of the principal curvatures of …
We give an explicit formula for the probability distribution based on a relativistic extension of Brownian motion. The distribution 1) is properly normalized and 2) obeys the tower law (semigroup property), so we can construct martingales and self-financing hedging strategies and price claims (options). This model is a…
Improved wind speed forecasts for power generation using machine learning.
Normalizing flows are a powerful class of generative models for continuous random variables, showing both strong model flexibility and the potential for non-autoregressive generation. These benefits are also desired when modeling discrete random variables such as text, but directly applying normalizing flows to discret…
Investigates market dynamics with informed traders and high-frequency traders.
A new path gradient estimator speeds up normalizing flows without sacrificing accuracy.
BatchNorm helps train quantized networks by avoiding gradient explosion.
The paper extends a Harnack inequality to noncompact evolving hypersurfaces.
Class Normalization improves zero-shot learning models.
We show that normalized currents of integration along the common zeros of random -tuples of sections of powers of singular Hermitian big line bundles on a compact Kähler manifold distribute asymptotically to the wedge product of the curvature currents of the metrics. If the Hermitian metrics are Hölder with sing…
The paper improves the empirical bootstrap method for non-normal estimators.
We consider a unit speed curve in Euclidean four-dimensional space and denote the Frenet frame by . We say that is a slant helix if its principal normal vector makes a constant angle with a fixed direction . In this work we give different characterizations of such curves in terms o…
Paper proposes energy objective for training normalizing flows without determinants.
Batch Normalization (BN) is a common technique used to speed-up and stabilize training. On the other hand, the learnable parameters of BN are commonly used in conditional Generative Adversarial Networks (cGANs) for representing class-specific information using conditional Batch Normalization (cBN). In this paper we pro…
The paper studies curvature flows of star-shaped hypersurfaces and proves convergence to spheres.
HollowFlow speeds up likelihood evaluation for large-scale models.
The paper studies a modified scalar curvature flow and proves convergence to a sphere.
Multiplicative stochasticity such as Dropout improves the robustness and generalizability of deep neural networks. Here, we further demonstrate that always-on multiplicative stochasticity combined with simple threshold neurons are sufficient operations for deep neural networks. We call such models Neural Sampling Machi…
A recent strategy to circumvent the exploding and vanishing gradient problem in RNNs, and to allow the stable propagation of signals over long time scales, is to constrain recurrent connectivity matrices to be orthogonal or unitary. This ensures eigenvalues with unit norm and thus stable dynamics and training. However …
Improved sampling efficiency for molecular systems using path gradients after Flow Matching.
Recently, we proposed to transform the outputs of each hidden neuron in a multi-layer perceptron network to have zero output and zero slope on average, and use separate shortcut connections to model the linear dependencies instead. We continue the work by firstly introducing a third transformation to normalize the scal…
A new method speeds up SoftMax normalization for embedding learning.
Driving styles have a great influence on vehicle fuel economy, active safety, and drivability. To recognize driving styles of path-tracking behaviors for different divers, a statistical pattern-recognition method is developed to deal with the uncertainty of driving styles or characteristics based on probability density…