Proof given for SGD convergence in a concise manner.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The adaptive moment estimation algorithm Adam (Kingma and Ba) is a popular optimizer in the training of deep neural networks. However, Reddi et al. have recently shown that the convergence proof of Adam is problematic and proposed a variant of Adam called AMSGrad as a fix. In this paper, we show that the convergence pr…
Proofs high-dimensional spectrum convergence of weighted sample covariance.
A common way to train neural networks is the Backpropagation. This algorithm includes a gradient descent method, which needs an adaptive step size. In the area of neural networks, the ADAM-Optimizer is one of the most popular adaptive step size methods. It was invented in \cite{Kingma.2015} by Kingma and Ba. The …
Proof of convergence for multi-objective optimization using inverse reinforcement learning.
AdaBoost's classifier and margins converge to a known value.
New proof shows faster convergence rate for robust estimation with Lasso in adversarially contaminated outputs.
In this article, we study the convergence of Mirror Descent (MD) and Optimistic Mirror Descent (OMD) for saddle point problems satisfying the notion of coherence as proposed in Mertikopoulos et al. We prove convergence of OMD with exact gradients for coherent saddle point problems, and show that monotone convergence on…
Adam converges to stationary points under relaxed conditions.
In this paper, we give an alternative proof for the convergence of Kähler-Ricci flow on a Fano mnaifold . This proof differs from that in [TZ3]. Moreover, we generalize the main theorem of [TZ3] to the case that may not admit any Kähler-Einstein metrics.
New proof shows local wealth condensation in economic models with biases.
New factorial power constants improve optimization convergence rates.
Paper proves convergence of measure transfer schemes using slicing and matching.
The paper proves stability of the positive mass theorem using intrinsic flat convergence.
New guarantees for black-box variational inference methods.
Quantile Temporal-Difference learning proved convergent with proof.
In this paper we will prove a maximum principle for the solutions of linear parabolic equation on complete non-compact manifolds with a time varying metric. We will prove the convergence of the Neumann Green function of the conjugate heat equation for the Ricci flow in to the minimal fundamental solut…
The Expectation-Maximization (EM) algorithm for mixture models often results in slow or invalid convergence. The popular convergence proof affirms that the likelihood increases with Q; Q is increasing in the M -step and non-decreasing in the E-step. The author found that (1) Q may and should decrease in some E-steps; (…
Direct proof shows adaptive gradient descent converges near-linearly for convex functions.
In this paper, as a study of reinforcement learning, we converge the Q function to unbounded rewards such as Gaussian distribution. From the central limit theorem, in some real-world applications it is natural to assume that rewards follow a Gaussian distribution , but existing proofs cannot guarantee convergence of th…
Gradient descent proves global convergence for deep networks with a single wide layer.
In this paper we study utility maximization with proportional transaction costs. Assuming extended weak convergence of the underlying processes we prove the convergence of the corresponding utility maximization problems. Moreover, we establish a limit theorem for the optimal trading strategies. The proofs are based on …
In this paper, we present a probability one convergence proof, under suitable conditions, of a certain class of actor-critic algorithms for finding approximate solutions to entropy-regularized MDPs using the machinery of stochastic approximation. To obtain this overall result, we prove the convergence of policy evaluat…
We prove convergence for suitably normalized solutions of the parabolic complex Monge-Ampère equation on compact Hermitian manifolds. This provides a parabolic proof of a recent result of Tosatti and Weinkove.
The multiplicative update (MU) algorithm has been extensively used to estimate the basis and coefficient matrices in nonnegative matrix factorization (NMF) problems under a wide range of divergences and regularizers. However, theoretical convergence guarantees have only been derived for a few special divergences withou…
Study shows convergence of anticanonically balanced metrics to Kähler-Einstein metrics on Fano manifolds.
We provide a simple proof of convergence covering both the Adam and Adagrad adaptive optimization algorithms when applied to smooth (possibly non-convex) objective functions with bounded gradients. We show that in expectation, the squared norm of the objective gradient averaged over the trajectory has an upper-bound wh…
This paper has been withdrawn by the author for further modification.
Proof of wall-crossing formula using spectral networks.
Paper proves MS convergence for radially symmetric kernels with large bandwidths.
We give the first rigorous proof of the convergence of Riemannian Hamiltonian Monte Carlo, a general (and practical) method for sampling Gibbs distributions. Our analysis shows that the rate of convergence is bounded in terms of natural smoothness parameters of an associated Riemannian manifold. We then apply the metho…
We study the evolution of anticanonical line bundles along the Kähler Ricci flow. We show that under some conditions, the convergence of Kähler Ricci flow is determined by the properties of the anticanonical divisors of . As examples, the Kähler Ricci flow on converges when is a Fano surface and …
The split Bregman (SB) method [T. Goldstein and S. Osher, SIAM J. Imaging Sci., 2 (2009), pp. 323-43] is a fast splitting-based algorithm that solves image reconstruction problems with general l1, e.g., total-variation (TV) and compressed sensing (CS), regularizations by introducing a single variable split to decouple …
Study shows long-term solutions for complex equations on curved spaces.
Deep networks converge in direction, with implications for predictions and margins.
We study some estimates along the Kahler Ricci flow on Fano manifolds. Using these estimates, we show the convergence of Kahler Ricci flow directly if the -invariant of the canonical class is greater than . Applying these convergence theorems, we can give a flow proof of Calabi conjecture on such Fano…
On certain del Pezzo surfaces with large automorphism groups, it is shown that the solution to the Kähler-Ricci flow with a certain initial value converges in -norm exponentially fast to a Kähler-Einstein metric. The proof is based on the method of multiplier ideal sheaves.
The paper analyzes stability and convergence rates of entropic and Sinkhorn potentials.
This note refers to our previous paper "The emergence of torsion in the continuum limit of distributed edge-dislocations". It identifies and fixes an error in the notion of convergence of Weitzenböck manifolds defined in the paper, and in the proof of the well-definiteness of this notion of convergence.
We give a new proof of the fact that the value function of the finite time horizon American put option for a jump diffusion, when the jumps are from a compound Poisson process, is the classical solution of a free boundary equation. We also show that the value function is across the optimal stopping boundary. Our …
The paper proves convergence of discrete maps to Riemann mappings for polyhedral surfaces.
Deep neural networks converge to Gaussian mixtures as layer width increases.
Paper proves convergence of SA algorithm via martingale and converse Lyapunov methods.
The paper studies a flow equation on even-dimensional manifolds, proving convergence under critical conditions.
The three operator splitting scheme was recently proposed by [Davis and Yin, 2015] as a method to optimize composite objective functions with one convex smooth term and two convex (possibly non-smooth) terms for which we have access to their proximity operator. In this short note we provide an alternative proof for the…
The paper studies invariant weighted Bergman metrics on domains.
New proof shows coupling-based flows converge linearly to diagonalize data covariance.
We provide a theoretical treatment of over-specified Gaussian mixtures of experts with covariate-free gating networks. We establish the convergence rates of the maximum likelihood estimation (MLE) for these models. Our proof technique is based on a novel notion of \emph{algebraic independence} of the expert functions. …