A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
We consider the case of derivative-free algorithms for non-convex optimization, also known as zero order algorithms, that use only function evaluations rather than gradients. For a wide variety of gradient approximators based on finite differences, we establish asymptotic convergence to second order stationary points u…
Given a 3-manifold M with no spherical boundary components, and a primitive class φin H^1(M;Z), we show that the following are equivalent: (1) φis a fibered class, (2) the rank gradient of (M,φ) is zero, (3) the Heegaard gradient of (M,φ) is zero.
The infimal Heegaard gradient of a compact 3-manifold was defined and studied by Marc Lackenby in an approach toward the well-known virtually Haken conjecture. As instructive examples, we consider Seifert fibered 3-manifolds, and show that a Seifert fibered 3-manifold has zero infimal Heegaard gradient if and only if i…
Uncertainty sampling, a popular active learning algorithm, is used to reduce the amount of data required to learn a classifier, but it has been observed in practice to converge to different parameters depending on the initialization and sometimes to even better parameters than standard training on all the data. In this…
Min-max formulations have attracted great attention in the ML community due to the rise of deep generative models and adversarial methods, while understanding the dynamics of gradient algorithms for solving such formulations has remained a grand challenge. As a first step, we restrict to bilinear zero-sum games and giv…
We show that gradient shrinking, expanding or steady Ricci solitons have potentials leading to suitable reference probability measures on the manifold. For shrinking solitons, as well as expanding soltions with nonnegative Ricci curvature, these reference measures satisfy sharp logarithmic Sobolev inequalities with low…
Overparameterized deep networks have the capacity to memorize training data with zero \emph{training error}. Even after memorization, the \emph{training loss} continues to approach zero, making the model overconfident and the test performance degraded. Since existing regularizers do not directly aim to avoid zero train…
We formulate a general framework for competitive gradient-based learning that encompasses a wide breadth of multi-agent learning algorithms, and analyze the limiting behavior of competitive gradient-based learning algorithms using dynamical systems theory. For both general-sum and potential games, we characterize a non…
It is well known that Markov chain Monte Carlo (MCMC) methods scale poorly with dataset size. A popular class of methods for solving this issue is stochastic gradient MCMC. These methods use a noisy estimate of the gradient of the log posterior, which reduces the per iteration computational cost of the algorithm. Despi…
We consider regular surfaces M that are given as the zeros of a polynomial function p:R3→R, where the gradient of p vanishes nowhere. We assume that M has non-zero mean curvature and prove that there exist only two examples of such surfaces, namely the sphere and the circular cylinder.
Graph manifolds are manifolds that decompose along tori into pieces with a tame S1-structure. In this paper, we prove that the simplicial volume of graph manifolds (which is known to be zero) can be approximated by integral simplicial volumes of their finite coverings. This gives a uniform proof of the vanishing of …
Recently, we proposed to transform the outputs of each hidden neuron in a multi-layer perceptron network to have zero output and zero slope on average, and use separate shortcut connections to model the linear dependencies instead. We continue the work by firstly introducing a third transformation to normalize the scal…
We prove that if the fundamental group of an orientable finite volume hyperbolic 3-manifold has finite index in the reflection group of a right-angled ideal polyhedra in H3 then it has a co-final tower of finite sheeted covers with positive rank gradient. The manifolds we provide are also known to have co-f…
New method trains neural networks with threshold activation functions efficiently.
problem Training neural networks with threshold activation functions is challenging due to zero gradients.
method We study weight decay regularized training problems of deep neural networks with threshold activations, showing they can be formulated as convex optimization problems.
result Regularized deep threshold network training problems can be formulated as standard convex optimization problems, paralleling the LASSO method.
We propose a simple and general variant of the standard reparameterized gradient estimator for the variational evidence lower bound. Specifically, we remove a part of the total derivative with respect to the variational parameters that corresponds to the score function. Removing this term produces an unbiased gradient …
In this paper, we show that steady or shrinking complete gradient Yamabe solitons with finite total scalar curvature and non-positive Ricci curvature are Ricci flat. Moreover, under certain pinching condition for Ricci curvature, we show that steady or shrinking complete gradient Yamabe solitons with finite total scala…
In this paper, we present some theoretical work to explain why simple gradient descent methods are so successful in solving non-convex optimization problems in learning large-scale neural networks (NN). After introducing a mathematical tool called canonical space, we have proved that the objective functions in learning…