Re-initializing neural networks improves generalization but not as much as other techniques.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New initialization techniques improve the performance and speed of EMI sensor-based object discrimination.
Single layer Feedforward Neural Network(FNN) is used many a time as a last layer in models such as seq2seq or could be a simple RNN network. The importance of such layer is to transform the output to our required dimensions. When it comes to weights and biases initialization, there is no such specific technique that co…
Paper proves new inequalities for Einstein-Maxwell data sets.
Paper proposes efficient tensor completion method using Gaussian Process.
We describe a proof of M.T. Anderson's result on the rigidity of complete stationary initial data for the Einstein vacuum equations in spacetime dimension 3 + 1, under an extra assumption on the norm of the stationary Killing vector field. The argument only involves basic comparison geometry along with some Bochner-Wei…
Deep neural networks (DNNs) form the backbone of almost every state-of-the-art technique in the fields such as computer vision, speech processing, and text analysis. The recent advances in computational technology have made the use of DNNs more practical. Despite the overwhelming performances by DNN and the advances in…
New method learns good initialization for gradient descent from past solutions.
White matter hyperintensity (WMH) is commonly found in elder individuals and appears to be associated with brain diseases. U-net is a convolutional network that has been widely used for biomedical image segmentation. Recently, U-net has been successfully applied to WMH segmentation. Random initialization is usally used…
Paper improves ISDA margin calculation using LSMC.
Spectral normalization stabilizes GANs by controlling gradient explosion and vanishing.
MIK improves t-SNE's local structure preservation in biological sequence data.
Improved neural network training by coupled initialization reduces neuron count.
Proposes a method to optimize neural network initialization using marginal likelihood maximization.
We establish a type of positive energy theorem for asymptotically anti-de Sitter Einstein-Maxwell initial data sets by using Witten's spinoral techniques.
Paper presents robust clustering methods for general mixture models.
In this paper, we present a novel approach for initializing deep neural networks, i.e., by turning PCA into neural layers. Usually, the initialization of the weights of a deep neural network is done in one of the three following ways: 1) with random values, 2) layer-wise, usually as Deep Belief Network or as auto-encod…
A new method for neural network initialization using graph degeneracy.
Initial margin requirements are becoming an increasingly common feature of derivative markets. However, while the valuation of derivatives under collateralisation (Piterbarg 2010, Piterbarg2012), under counterparty risk with unsecured funding costs (FVA) (Burgard2011, Burgard2011, Burgard2013) and in the presence of re…
Paper uses Chebyshev Tensors for accurate dynamic sensitivities and ISDA SIMM computation.
New pruning methods improve energy efficiency of neural networks.
The paper tackles fVaR prediction methods in finance.
Many modern learning tasks involve fitting nonlinear models to data which are trained in an overparameterized regime where the parameters of the model exceed the size of the training dataset. Due to this overparameterization, the training loss may have infinitely many global minima and it is critical to understand the …
Mimetic initialization improves Transformer training on small datasets.
Improved LLM pre-training performance through better weight and variance control.
The techniques developed by Butscher in arXiv:math/0703469 for constructing constant mean curvature (CMC) hypersurfaces in the (n+1)-sphere by gluing together spherical building blocks are generalized to handle less symmetric initial configurations. The outcome is that the approximately CMC hypersurface obtained by glu…
This work investigates the ways in which deep learning methods can benefit from random projection (RP), a classic linear dimensionality reduction method. We focus on two areas where, as we have found, employing RP techniques can improve deep models: training neural networks on high-dimensional data and initialization o…
Study peels tensor equations on Schwarzschild spacetime.
This is the first in a series Of papers in which we initiate the study Of very rough solutions to the initial value problem for the Einstein Vacuum equations expressed relative to wave coordinates. By very rough we mean solutions which cannot be constructed by the classical techniques Of energy estimates and Sobolev in…
Proposes a novel approach for deep neural network initialization using polynomial approximations.
K-Means clustering improved with sophisticated initialisation techniques.
We show that there exists a suitable neighborhood of a constant curvature hyperbolic metric such that, for all initial data in this neighborhood, the corresponding solution to a normalized cross curvature flow exists for all time and converges to a hyperbolic metric. We show that the same technique proves an analogous …
This article, written to appear as a chapter in "The Springer Handbook of Spacetime", is a review of the initial value problem for Einstein's gravitational field theory in general relativity. Designed to be accessible to graduate students who have taken a first course in general relativity, the article first discusses …
New technique trains deep neural networks without normalization or minibatch statistics.
Improves deep neural network training and accuracy with adaptive basis approach.
We provide techniques for studying the nonnegatively curved left-invariant metrics on a compact Lie group. For "straight" paths of left-invariant metrics starting at bi-invariant metrics and ending at nonnegatively curved metrics, we deduce a nonnegativity property of the initial derivative of curvature. We apply this …
Trained neural networks perform Bayesian reasoning for tasks beyond their initial scope.
Paper proves rigidity for spin bands with specific conditions.
This is the second in a series of three papers in which we initiate the study of very rough solutions to the initial value problem for the Einstein vacuum equations expressed relative to wave coordinates. By very rough we mean solutions which cannot be constructed by the classical techniques of energy estimates and Sob…
New method initializes MLPs for tabular data with tree-based feature interactions.
This paper studies how neural network architecture affects the speed of training. We introduce a simple concept called gradient confusion to help formally analyze this. When gradient confusion is high, stochastic gradients produced by different data samples may be negatively correlated, slowing down convergence. But wh…
Nonconvex matrix recovery is known to contain no spurious local minima under a restricted isometry property (RIP) with a sufficiently small RIP constant . If is too large, however, then counterexamples containing spurious local minima are known to exist. In this paper, we introduce a proof technique that is capa…
Training recurrent neural networks (RNNs) on long sequence tasks is plagued with difficulties arising from the exponential explosion or vanishing of signals as they propagate forward or backward through the network. Many techniques have been proposed to ameliorate these issues, including various algorithmic and archite…
New method improves BO's AF maximizer initialization for high-dimensional problems.
A new design methodology for neural networks that is guided by traditional algorithm design is presented. To prove our point, we present two heuristics and demonstrate an algorithmic technique for incorporating additional weights in their signal-flow graphs. We show that with training the performance of these networks …
We use numerical techniques to study the formation of singularities in Ricci flow. Comparing the Ricci flows corresponding to a one parameter family of initial geometries on S^3 with varying amounts of S^2 neck pinching, we find critical behavior at the threshold of singularity formation.
Gradient descent proves global convergence for 4-layer matrix factorization.
In this paper, we study the problem of expected utility maximization of an agent who, in addition to an initial capital, receives random endowments at maturity. Contrary to previous studies, we treat as the variables of the optimization problem not only the initial capital but also the number of units of the random end…