A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
We consider the case of derivative-free algorithms for non-convex optimization, also known as zero order algorithms, that use only function evaluations rather than gradients. For a wide variety of gradient approximators based on finite differences, we establish asymptotic convergence to second order stationary points u…
Given a 3-manifold M with no spherical boundary components, and a primitive class φin H^1(M;Z), we show that the following are equivalent: (1) φis a fibered class, (2) the rank gradient of (M,φ) is zero, (3) the Heegaard gradient of (M,φ) is zero.
The infimal Heegaard gradient of a compact 3-manifold was defined and studied by Marc Lackenby in an approach toward the well-known virtually Haken conjecture. As instructive examples, we consider Seifert fibered 3-manifolds, and show that a Seifert fibered 3-manifold has zero infimal Heegaard gradient if and only if i…
Uncertainty sampling, a popular active learning algorithm, is used to reduce the amount of data required to learn a classifier, but it has been observed in practice to converge to different parameters depending on the initialization and sometimes to even better parameters than standard training on all the data. In this…
Min-max formulations have attracted great attention in the ML community due to the rise of deep generative models and adversarial methods, while understanding the dynamics of gradient algorithms for solving such formulations has remained a grand challenge. As a first step, we restrict to bilinear zero-sum games and giv…
We show that gradient shrinking, expanding or steady Ricci solitons have potentials leading to suitable reference probability measures on the manifold. For shrinking solitons, as well as expanding soltions with nonnegative Ricci curvature, these reference measures satisfy sharp logarithmic Sobolev inequalities with low…
We formulate a general framework for competitive gradient-based learning that encompasses a wide breadth of multi-agent learning algorithms, and analyze the limiting behavior of competitive gradient-based learning algorithms using dynamical systems theory. For both general-sum and potential games, we characterize a non…
It is well known that Markov chain Monte Carlo (MCMC) methods scale poorly with dataset size. A popular class of methods for solving this issue is stochastic gradient MCMC. These methods use a noisy estimate of the gradient of the log posterior, which reduces the per iteration computational cost of the algorithm. Despi…
We consider regular surfaces M that are given as the zeros of a polynomial function p:R3→R, where the gradient of p vanishes nowhere. We assume that M has non-zero mean curvature and prove that there exist only two examples of such surfaces, namely the sphere and the circular cylinder.
Graph manifolds are manifolds that decompose along tori into pieces with a tame S1-structure. In this paper, we prove that the simplicial volume of graph manifolds (which is known to be zero) can be approximated by integral simplicial volumes of their finite coverings. This gives a uniform proof of the vanishing of …
Recently, we proposed to transform the outputs of each hidden neuron in a multi-layer perceptron network to have zero output and zero slope on average, and use separate shortcut connections to model the linear dependencies instead. We continue the work by firstly introducing a third transformation to normalize the scal…
We prove that if the fundamental group of an orientable finite volume hyperbolic 3-manifold has finite index in the reflection group of a right-angled ideal polyhedra in H3 then it has a co-final tower of finite sheeted covers with positive rank gradient. The manifolds we provide are also known to have co-f…
New method trains neural networks with threshold activation functions efficiently.
problem Training neural networks with threshold activation functions is challenging due to zero gradients.
method We study weight decay regularized training problems of deep neural networks with threshold activations, showing they can be formulated as convex optimization problems.
result Regularized deep threshold network training problems can be formulated as standard convex optimization problems, paralleling the LASSO method.
We propose a simple and general variant of the standard reparameterized gradient estimator for the variational evidence lower bound. Specifically, we remove a part of the total derivative with respect to the variational parameters that corresponds to the score function. Removing this term produces an unbiased gradient …