SGD fails to converge for deep ReLU networks with limited random initializations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Stochastic variational inference is an established way to carry out approximate Bayesian inference for deep models. While there have been effective proposals for good initializations for loss minimization in deep learning, far less attention has been devoted to the issue of initialization of stochastic variational infe…
New MKABSDEs help calculate initial margins in financial contracts.
New insights show stochastic initialization prevents token clustering in deep Transformers.
We study the problem of training deep neural networks with Rectified Linear Unit (ReLU) activation function using gradient descent and stochastic gradient descent. In particular, we study the binary classification problem and show that for a broad family of loss functions, with proper random weight initialization, both…
Improves state space models' resistance to noise.
In recent literature, a general two step procedure has been formulated for solving the problem of phase retrieval. First, a spectral technique is used to obtain a constant-error initial estimate, following which, the estimate is refined to arbitrary precision by first-order optimization of a non-convex loss function. N…
Study on the smoothness of solutions to a specific type of stochastic differential equation.
Wave maps with noise can lead to self-similar blowup from arbitrary initial data.
We study the convergence properties of the VR-PCA algorithm introduced by \cite{shamir2015stochastic} for fast computation of leading singular vectors. We prove several new results, including a formal analysis of a block version of the algorithm, and convergence from random initialization. We also make a few observatio…
Many modern learning tasks involve fitting nonlinear models to data which are trained in an overparameterized regime where the parameters of the model exceed the size of the training dataset. Due to this overparameterization, the training loss may have infinitely many global minima and it is critical to understand the …
New method initializes MLPs for tabular data with tree-based feature interactions.
The paper generalizes the construction by stochastic flows of consistent utility processes introduced by M. Mrad and N. El Karoui in (2010). The utilities random fields are defined from a general class of processes denoted by $\GX$. Making minimal assumptions and convex constraints on test-processes, we construct by co…
Enlargement of filtrations is a classical topic in the general theory of stochastic processes. This theory has been applied to stochastic finance in order to analyze models with insider information. In this paper we study initial enlargement in a Markov chain market model, introduced by R. Norberg. In the enlargened fi…
Study optimizes trading in multiple assets with cross-effects.
We provide sufficient conditions on the coefficients of a stochastic evolution equation on a Hilbert space of functions driven by a cylindrical Wiener process ensuring that its mild solution is positive if the initial datum is positive. As an application, we discuss the positivity of forward rates in the Heath-Jarrow-M…
SGD transitions between maxima and minima with varying time scales.
Study on reproducibility in optimization with bounds on limits.
The goal of this paper is to clarify when a stochastic partial differential equation with an affine realization admits affine state processes. This includes a characterization of the set of initial points of the realization. Several examples, as the HJMM equation from mathematical finance, illustrate our results.
We analyze single-layer neural networks with the Xavier initialization in the asymptotic regime of large numbers of hidden units and large numbers of stochastic gradient descent training steps. The evolution of the neural network during training can be viewed as a stochastic system and, using techniques from stochastic…
Determining the 3D structures of biological molecules is a key problem for both biology and medicine. Electron Cryomicroscopy (Cryo-EM) is a promising technique for structure estimation which relies heavily on computational methods to reconstruct 3D structures from 2D images. This paper introduces the challenging Cryo-…
In this paper, we study the minimax optimization problem in the smooth and strongly convex-strongly concave setting when we have access to noisy estimates of gradients. In particular, we first analyze the stochastic Gradient Descent Ascent (GDA) method with constant stepsize, and show that it converges to a neighborhoo…
Deep neural networks achieve state-of-the-art performance for a range of classification and inference tasks. However, the use of stochastic gradient descent combined with the nonconvexity of the underlying optimization problems renders parameter learning susceptible to initialization. To address this issue, a variety o…
Safe learning of stochastic dynamics with safety constraints.
New initialization schemes preserve fractional moments of weights in deep networks, improving training and test performance.
Posterior inference in directed graphical models is commonly done using a probabilistic encoder (a.k.a inference model) conditioned on the input. Often this inference model is trained jointly with the probabilistic decoder (a.k.a generator model). If probabilistic encoder encounters complexities during training (e.g. s…
Why does training deep neural networks using stochastic gradient descent (SGD) result in a generalization error that does not worsen with the number of parameters in the network? To answer this question, we advocate a notion of effective model capacity that is dependent on {\em a given random initialization of the netw…
This paper presents a stochastic behavior analysis of a kernel-based stochastic restricted-gradient descent method. The restricted gradient gives a steepest ascent direction within the so-called dictionary subspace. The analysis provides the transient and steady state performance in the mean squared error criterion. It…
Gradient descent training of neural networks leads to solutions close to natural cubic splines.
Investment strategies in occupational pension plans are optimized for non-tradable income risk.
CDS combines PT and diffusion for efficient sampling from multimodal distributions.
We present an approach for efficiently training Gaussian Mixture Model (GMM) by Stochastic Gradient Descent (SGD) with non-stationary, high-dimensional streaming data. Our training scheme does not require data-driven parameter initialization (e.g., k-means) and can thus be trained based on a random initialization. Furt…
New analysis shows temperature guarantees generalization in stochastic training.
Persistent neurons improve neural network optimization by leveraging previous solutions.
Deep Hedging learns optimal strategies for various risk levels.
Stochastic gradient descent outperforms traditional force-directed methods.
New method uses SDEs for accurate non-uniformly sampled time series analysis.
One of the difficulties of training deep neural networks is caused by improper scaling between layers. Scaling issues introduce exploding / gradient problems, and have typically been addressed by careful scale-preserving initialization. We investigate the value of preserving scale, or isometry, beyond the initial weigh…
New analysis improves convergence guarantees for diffusion-based samplers in Wasserstein distance.
Optimal algorithm for high-dimensional stochastic linear bandits with sparse parameters.
We solve a complex trade execution problem by simplifying it into a known LQ control problem.
Deep networks retain initial bias after training, affecting generalization.
Paper shows pre-training and transfer learning reduce sample complexity for neural networks.
This paper examines the problem of pricing spread options under some models with jumps driven by Compound Poisson Processes and stochastic volatilities in the form of Cox-Ingersoll-Ross(CIR) processes. We derive the characteristic function for two market models featuring joint normally distributed jumps, stochastic vol…
We propose a new high-order alternating direction implicit (ADI) finite difference scheme for the solution of initial-boundary value problems of convection-diffusion type with mixed derivatives and non-constant coefficients, as they arise from stochastic volatility models in option pricing. Our approach combines differ…
Stochastic gradient descent converges to universal limits in high dimensions.
We show that stochastic interpolation flow maps are Lipschitz with a sharp constant.
Paper introduces stochastic HJB on Jacobi structures.